About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore’s AI infrastructure platforms. This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters. The Team Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. The Systems Engineering and Platform Validation team ensures Graphcore’s AI compute platforms are reliable, diagnosable, and operationally robust at scale. The team co
Jobs in United States
Principal Hardware Diagnostics Engineer in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current principal hardware diagnostics engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced Silicon Test Engineer to join our Product Test and Diagnosis Department (PTD). This is a pivotal role and will involve building a team of engineers to develop System Level Test (SLT) capability within the company. Working closely with a cross-functional team you will implement SLT tests for a family of next generation AI Processors. The ideal candidate should have a focus on quality and demonstrate a good understanding of the importance of production test on the success of a product. T hey will have a proven Functional Test or ATE Test Engineering background, and will have a pragmatic, hands-on and flexible approach to a fast-changing environment. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Managing a team of
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our Tensix Team is building the future of AI compute with a ground-up architecture centered on scalable RISC-V processors. As we push performance boundaries, we’re reimagining the frontend of our RISC-V cores to deliver major gains in programmability, efficiency, and developer experience. This is a rare opportunity to shape the CPU architecture at the heart of our AI platform and lead one of the most strategic technical efforts at Tenstorrent. This role is hybrid, based out of Toronto, ON, Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced Microarchitect: 10+ years of deep expertise in CPU performance modeling and microarchitecture design. AI Workload Expert: Deeply familiar with the computational and memory bottlenecks of modern AI workloads, particularly Large Language Models (LLMs). Hardware-Software Co-Designer: Driven to architect custom instruction set extensions and validate their performance gains against real-world workloads. Ways to stand-out: Familiarity with open-source RISC-V cores, AI-based agentic workflow experience What We Need Profile & Analyze: Dissect cutting-edge AI workloads to identi
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad
We are hiring senior engineers to work on the CUDA driver, a core component of our platform for accelerating general purpose computation on the GPU. Our team delivers features and improvements to better realize the potential of NVIDIA hardware for a growing range of computational workloads, ranging from deep learning, scientific computation, and self-driving cars to video games and virtual reality! CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, switches and the operating system. What you'll be doing: As a member of our team, you will use your design abilities, coding expertise, and creativity to deliver the best Compute platform in the world. You will craft elegant solutions to exciting problems and craft the future direction of CUDA as you collaborate with your peers across NVIDIA. You will evangelize, architect, and implement new CUDA features You'll oversee and drive development efforts across multiple teams Collaborate with members of hardware architecture teams Help define forward-looking improvements to the CUDA APIs and programming model Design and maintain performance and precision modeling Write effective, maintainable, and well-tested code Develop code for multiple operating systems What we need to see: Bachelor of Science or Master of Science degree in Computer Science, Electrical Engineering, or related field (or equivalent experience) 15+ years of relevant systems software development experience Strong C programming skills </
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects. Job Summary We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines. The Team The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers. The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale. As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems
About us Graphcore is one of the world’s leading innovators in artificial intelligence compute. We are developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and support the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a family of companies responsible for some of the world’s most transformative technologies. Together, we share a bold vision to enable advanced artificial intelligence and ensure its benefits are accessible to everyone. Graphcore brings together AI researchers, silicon designers, software engineers and systems architects to solve complex technical challenges and deliver innovative computing solutions. Job Summary The Principal Electrical Engineer will be a technical authority within Data Center Engineering, leading the architecture and delivery of safe, resilient and scalable electrical infrastructure for high-density AI computing environments. Working with internal teams, data center developers, utilities, consultants and equipment partners, this role will guide projects from early technical studies through design, construction, commissioning, operation and lifecycle improvement. The successful candidate must reside in, or be willing to relocate to, Austin, Texas. Approximately 10% travel may be required. The Team The Data Center Engineering team is responsible for defining and enabling the infrastructure needed to deploy and operate Graphcore’s computing systems at scale. The team works across electrical, mechanical, thermal, controls, systems and operational disciplines, collaborating with external engineering and construction partners to deliver reliable, efficient and maintainable data center environments. Responsibilities and Duties Act as the technical authority for electrical engineering across data center infrastructure projects, from the utility or on-site power source through to the IT rack. Lead electrical archit
Work Flexibility: Hybrid or Onsite Stryker is seeking a dynamic and visionary product development leader with a proven track record of leading large, complex product development programs through the application of systems engineering principles. As the technical quarterback, this individual will drive alignment, collaboration, and execution across a multidisciplinary global product line, ensuring teams work cohesively toward shared objectives while delivering innovative solutions that improve patient care. As medical technologies continue to evolve, our products must operate seamlessly within broader healthcare ecosystems and enabling technology platforms. This leader will be responsible for defining and overseeing system-level architectures, ensuring robust integration across hardware, software, data, connectivity, and partner technologies. They will champion a systems-thinking approach, proactively identify opportunities and risks, and ensuring our products deliver a seamless user experience while supporting the future needs of a connected healthcare environment. The ideal candidate combines deep technical expertise with exceptional leadership, influencing stakeholders across functions, geographies, and organizational levels to bring complex systems and product visions to life. https://www.stryker.com/us/en/nse.html What You Will Do Lead system architecture, technical strategy, and integration activities for complex medical device products and platforms. Translate customer, clinical, and business needs into system-level requirements and drive alignment across software, electrical, and mechanical engineering teams. Lead requirements development, allocation, traceability, and verification planning thro
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Firmware & Product Test (FPT) team plays a critical role in delivering high-quality enterprise SSD solutions by ensuring firmware functionality, reliability, and compliance. We work across simulation, FPGA, and hardware environments to validate modern storage technologies, build scalable automation, and drive continuous improvement in validation methodologies. Our team values technical excellence, collaboration, and innovation, including the use of AI-enabled tools to enhance engineering productivity and quality. As a Principal Test Development Engineer, you will serve as a technical leader for firmware validation, defining verification strategies, advancing automation frameworks, and driving complex failure analysis efforts. This role offers the opportunity to influence product quality across multiple SSD programs while mentoring engineers and partnering closely with firmware architects to improve testability and validation effectiveness. Responsibilities: Lead verification strategy, test planning, automation, and coverage closure for NVMe front-end firmware features across multiple product lines Architect and enhance scalable Python-based test automation frameworks, CI/CD integration, regression infrastructure, and reporting capabilities Drive root-cause analysis and failure triage using firmware traces, protocol analyzers, system logs, and structured debug methodologies Define validation standards, review test code, mentor engineers, and promote standard methodologies in automation and qua
$147.1K – $230.9K/yr
Principal Embedded Firmware and Software Engineer Description - We are seeking a Principal Embedded Firmware & Software Engineer to lead the design, development, and debugging of embedded software and firmware for computer systems. In this role, you will combine deep, hands-on engineering expertise with system-level technical leadership to ensure seamless integration between software and hardware components, delivering reliable and efficient system performance. You will collaborate closely with cross-functional teams including hardware engineers, software developers, QA, and product managers to bring high-quality products to market. Responsibilities Provide technical leadership for the architecture, development, security, integration, debugging, validation, and deployment of embedded firmware and software including BIOS/UEFI, EFI applications and drivers, embedded controllers, and RTOS-based systems. Analyze hardware and system architectures to define firmware requirements, dependencies, interfaces, integration strategies, and validation approaches. Troubleshoot and resolve firmware issues by designing and implementing enhancements, updates, and programming changes across firmware subsystems. Define and drive firmware integration, verification, and validation strategies, including automated testing, regression testing, and continuous integration . Develop and improve engineering tools and automation using Python and other appropriate technologies for development, debugging, testing, analysis, and validation. Advance CI/CD and DevSecOps practices for embedded development to improve engineering velocity, quality, traceability, and release confidence. Evaluate and apply AI-assisted software development and engineering tools where they can improve developer productivity,
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. About the Team Cloudflare handles traffic for almost 25% of the Internet. That’s a lot of data. On the Town Lake team, our mission is to make that data accessible and valuable for users across the company. We connect data from dozens of source systems and make it available so that any user in the company can answer any question in 5 minutes or less, using SQL or plain english. We’re building a modern, agentic-first data lakehouse platform ba
About the Team Tax and Trade at OpenAI shapes business strategy by embedding critical tax, export control, customs, and cross-border considerations into how the company builds, sources, scales, and operates in support of the mission. We combine deep expertise with practical systems thinking to look around corners, identify emerging risks and opportunities early, and help teams make smarter decisions at the point where strategy becomes execution. Across procurement, hardware operations, manufacturing, logistics, finance, legal, supplier onboarding, and operator workflows, we build robust, scalable support services leveraging cutting-edge technology—including governed AI and automation—to make complex regulated work more durable, more efficient, and easier to scale. About the Role We’re hiring a Senior Manager, Export Controls to lead OpenAI’s export controls strategy and operating model. This is a senior role with broad scope across advanced computing, semiconductors, software, hardware, manufacturing, and high technology partnerships. You will refine how OpenAI classifies controlled technology, software, and hardware, structures access-controlled environments, manages licensing and supplier commitments, and scales export-control operations in a way that supports the company’s pace of innovation. You will also shape how OpenAI applies AI and agentic workflows to policy-heavy operational work, building systems that make complex rules easier to navigate and easier to execute. In this role, you will: Refine the strategy and operating model for OpenAI’s export controls program across advanced computing, semiconductors, software, hardware, manufacturing, and high technology partnerships. Own export classification and licensing strategy for controlled technical data, software, hardware, and research environments. Lead the design and operation of compliant controlled environments and related governance processes. Partner with Research and Infrastructure to support efficient
This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. This is where your work saves lives As a Principal Android Software Engineer, you will lead architecture and delivery of shared Android capabilities that power multiple clinical product applications worldwide. You will own significant portions of the Android system design—reusable services and libraries, hybrid WebView/React UI foundations, device connectivity patterns, and common clinical workflows—while mentoring engineers and raising the bar for quality in a regulated medical-device environment. An ideal candidate brings deep hands-on Android expertise, a clear technical vision for multi-product platform software, and the ability to drive complex cross-team decisions independently. Strong communication and collaboration with product teams, systems, hardware, and other platform partners are essential. What you'll be doing: Lead Platform Android Architecture and Delivery: Define and evolve shared Android platform architecture (Kotlin services, DI, WebView/JS bridge patterns, messaging/broker integr
Other cities to consider
More places hiring for this role
Get new principal hardware diagnostics engineer jobs in United States by email
Daily job updates · Unsubscribe anytime