About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Operating Systems team is critical in this mission, turning sophisticated hardware and AI capabilities into a reliable, trusted platform. Security is central to that work: we define trust boundaries, integrate hardware-backed protections, isolate sensitive context, and establish the guardrails that let AI applications and agents act safely, privately, and under user control. About the Role We’re looking for a Software Security Architect to define the security architecture for OpenAI’s next-generation operating system. You’ll work alongside hardware security architects and partner with operating system, silicon, firmware, privacy, and product teams to protect users, their devices, and their data. This is a senior, hands-on role for someone who can connect operating system internals, hardware-backed security, and real-world product constraints. Your work will shape the platform’s trust model, protect sensitive information, and set the technical foundations for AI that is safe, private, and under user control. In this role, you will: Define the operating system’s security architecture, trust boundaries, privilege model, and protections for sensitive user data. Partner with hardware security architects to integrate roots of trust, secure elements, trusted execution environments, and processor security capabilities into the operating system. Desig
Jobs in United States
Hardware Lead Engineer in United States
400 active opportunities · Updated October 2026
Showing
15 jobs
Explore current hardware lead engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Firmware & Product Test (FPT) team plays a critical role in delivering high-quality enterprise SSD solutions by ensuring firmware functionality, reliability, and compliance. We work across simulation, FPGA, and hardware environments to validate modern storage technologies, build scalable automation, and drive continuous improvement in validation methodologies. Our team values technical excellence, collaboration, and innovation, including the use of AI-enabled tools to enhance engineering productivity and quality. As a Principal Test Development Engineer, you will serve as a technical leader for firmware validation, defining verification strategies, advancing automation frameworks, and driving complex failure analysis efforts. This role offers the opportunity to influence product quality across multiple SSD programs while mentoring engineers and partnering closely with firmware architects to improve testability and validation effectiveness. Responsibilities: Lead verification strategy, test planning, automation, and coverage closure for NVMe front-end firmware features across multiple product lines Architect and enhance scalable Python-based test automation frameworks, CI/CD integration, regression infrastructure, and reporting capabilities Drive root-cause analysis and failure triage using firmware traces, protocol analyzers, system logs, and structured debug methodologies Define validation standards, review test code, mentor engineers, and promote standard methodologies in automation and qua
NVIDIA is seeking a strong technology leader to manage our Server Software Technical Program Management (TPM) team. This role is at the cross-section of execution and strategy, leading a team of Senior TPMs who drive the firmware and system software for NVIDIA's next-generation server platforms like DGX, MGX, and HGX. These platforms bring together the full power of NVIDIA GPUs, NVLink, InfiniBand networking, Grace CPUs, and our optimized AI/HPC software stack. This deep technical leadership role focused on the Software Development Processes that brings new server hardware to life. What you'll be doing: Lead a team of TPMs driving the technical software and firmware execution for NVIDIA's NPI (New Product Introduction) and sustaining engineering teams. Drive the end-to-end SDLC for low-level server components, including firmware (BMC, UEFI/BIOS), drivers, and system management software, ensuring alignment with hardware schedules. Collaborate closely with NVIDIA product management and hardware engineering teams to define release plans and program objectives. Build a strong connection and feedback loop between sustaining and NPI engineering teams to improve product quality and development velocity. Lead process improvement initiatives and help propagate SDLC standards across multiple engineering and TPM organizations. You will have the opportunity to interact with diverse technical groups, spanning all organizational levels. What we need to see: Bachelor of Science (or equivalent experience) or Master of Science degree in Computer Science, Electrical Engineering, or related field. 12+ overall years of experience developing and leading complex low-level or system software projects. and 7+ years of experience in a people management role. Deep understanding of system a
NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&
About the Team The Future of Computing Research team is an applied research team within OpenAI’s Consumer Devices group. We study how AI systems perceive people and their surroundings, and we turn that research into capabilities for future products. Our work spans machine learning, sensing, and hardware, with a focus on building systems that work beyond controlled environments. About the Role We’re looking for a machine learning engineer to help shape how future AI systems understand the physical world and the people in it. The role focuses on multimodal perception and authentication, bringing together signals from cameras, microphones, and other sensors. You’ll work with specialized perception models and larger multimodal models, and partner with hardware, firmware, software, and product teams to bring new research into real-world systems. This role is based in San Francisco. We work in the office three days per week and offer relocation assistance. In this role, you will: Research and develop multimodal perception and authentication methods across visual, audio, and other sensing signals. Explore how specialized perception models and larger multimodal models can work together. Design data, training, and evaluation approaches that improve performance in real-world conditions. Study model behavior, robustness, and failure modes across sensing, data, and deployment environments. Integrate and validate new capabilities in real-time or resource-constrained systems. Work with hardware, firmware, software, and product teams to turn research into working systems. You might thrive in this role if you: Have a strong background in computer vision, audio or speech machine learning, multimodal learning, or sensing. Have experience developing specialized machine learning models, larger multimodal models, or both. Have brought research ideas into practical systems, prototypes, or products. Know how to design experiments, build evaluations, and investigate model behavior. Have wo
About the Team Our mission at OpenAI is to discover and enact the path to safe, beneficial AGI. To do this, we believe that many technical breakthroughs are needed in generative modeling, reinforcement learning, large-scale optimization, active learning, and other areas. The team builds the performance-critical systems that allow OpenAI's models to run efficiently across a diverse set of AI accelerators. We work across the inference stack, from low-level kernels and compilers through model execution, to unlock the full capabilities of the underlying hardware. About the Role As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack required to run cutting-edge frontier models efficiently on the platform. This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution. You will work on the systems needed to support OpenAI's inference stack on Trainium, including developing and optimizing high-performance kernels, improving compiler support, and enabling efficient execution of the model forward pass. You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take full advantage of Trainium. The work may range from low-level hardware-specific optimization to compiler and runtime improvements to integrating new model architectures into the inference stack. If you enjoy working at the intersection of ML systems, compilers, kernels, and accelerator hardware, this role is for you. We're looking for engineers who are self-directed, comfortable operating across abstraction layers, and excited to solve challenging performance problems for frontier-scale AI systems. In This Role, You Will Build and optimize OpenAI's inference stack for AWS Trainium. Develop high-performance kernels for critical model operations and workloads. Extend and improve compiler support to efficiently target
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b
From $79.5K/yr
Location Details: Santa Clara, CA At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is an in-office position and you’ll be expected to work full-time in an office location, and therefore must live within commuting distance from your assigned office. You will work from this office beginning on your first day. Join our team Join GoDaddy's Global IT Support team as a Desktop Support Technician in our Santa Clara office, where you'll be the face of IT for employees who rely on technology to do their best work every day. In this hands-on, 100% on-site role, you'll provide walk-up and escalated technical support, troubleshooting hardware, software, workplace technology, and connectivity issues while delivering an exceptional customer experience.You'll collaborate closely with a globally distributed team of support professionals across North America, EMEA, and APAC, helping resolve both everyday technical requests and high-priority incidents in a fast-paced environment. Success in this role is driven as much by your communication, empathy, and problem-solving skills as your technical knowledge, making it an excellent opportunity for someone who enjoys helping people and learning new technologies.If you're passionate about customer service, curious about how technology works, and looking to grow your career in enterprise IT, we'd love to hear from you. What you'll get to do... Provide front-line technical support to employees by diagnosing and resolving hardware, software, operating system, peripheral, and account-related issues, ensuring minimal disruption to productivity. Manage and prioritize incidents and service requests through the IT ticketing system, taking ownership from initial intake through troubleshooting, resolution, and follow-up communication. Configure, deploy, maintain,
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. About the Role The Product Platform Tools team builds and operates Cloudflare's internal support and admin platform - the foundation that customer-facing and operational teams across the company rely on to do their jobs. As a Systems Engineer on this team, you'll design and build the backend services, infrastructure, APIs, and integrations that power this platform at enterprise scale, working closely with both engineering teams and non-engineering
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. About the Team Cloudflare handles traffic for almost 25% of the Internet. That’s a lot of data. On the Town Lake team, our mission is to make that data accessible and valuable for users across the company. We connect data from dozens of source systems and make it available so that any user in the company can answer any question in 5 minutes or less, using SQL or plain english. We’re building a modern, agentic-first data lakehouse platform ba
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations Austin, US About the Role Cloudflare's People team supports 5,000+ employees globally. To scale, we are building an AI-driven operating layer on the Cloudflare Developer Platform to automate workflows, ensure data integrity, and streamline employee support. You will ship production systems for hiring, onboarding, and self-service, using AI to create leverage while designing rigorous guardrails for sensitive employee data. Lever
About the Team At OpenAI, the User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving company priorities. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are looking for a senior program manager to build the safety, quality, and risk operations supporting a new category of consumer devices. You will translate ambiguous product risks and evolving requirements into practical operating models, workflows, escalation paths, launch-readiness plans, and cross-functional decision-making. This is a foundational role: the systems you build will shape how OpenAI launches, monitors, and improves a new category of consumer devices safely at scale. You will help establish how potential safety incidents, product-quality concerns, sensitive customer escalations, privacy-sensitive issues, and other emerging device risks are identified, investigated, resolved, and incorporated into product and operational improvements. You will turn incomplete requirements into practical workflows, decision rights, launch plans, quality controls, measurement, and durable ownership. The role centers on program building, operational judgment, and execution. We welcome candidates from product safety, quality assurance, regulatory operations, technical program management, healthcare, medical devices, aerospace, consumer technology, and other environments involving complex products or regulated risks. Direct consumer-hardware experience is helpful but not required. The strongest candidates learn unfam
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o
Other cities to consider
More places hiring for this role
Get new hardware lead engineer jobs in United States by email
Daily job updates · Unsubscribe anytime