Jobs in United States

Hardware Operations Engineer in United States

408 active opportunities · Updated October 2026

Explore current hardware operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an

KubernetesGitAIGo
M
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Own and evolve the Quality Management System (QMS) to support a regulated medical device development program, including design controls and DHF maintenance. Establish and enforce requirements traceability: user needs → design requirements → verification/validation artifacts and change control. Define and run the program-level V&V strategy (verification, validation, and test coverage), including test plans, protocols, reports, and acceptance criteria. Drive risk management activities (e.g., DFMEA / PFMEA, hazard analyses) and ensure mitigations are reflected in requirements and verification. Lead document control: reviews, approvals, training, retention, and audit readiness. Partner with engineering to make quality “native” to the dev workflow (automated testing, release gates, software configuration management). Prepare the program for audits and inspections, including hands-on audit leadership. What we’re looking for Senior experience leading quality for complex hardware + software products in a regulated environment. Deep familiarity with design controls, DHF, document control, risk management, and verification planning. Strong systems thinking and the ability to translate ambiguous product intent into testable requirements. Comfortable collaborating directly with multidisciplinary engineering (recon/ML, embedded, mechanical, EE, cloud). Useful experience Regulated product quality leadership (ISO 13485 / 21 CFR 820 or equivalent), including audit readiness and FDA-facing work. eQMS + document control fluency (e.g., Greenlight Guru) that integrates cleanly with modern engineering workflows.

M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Be the generalist EE for the scanner system: integration, bring-up, debugging, and making the electrical side of the device reliable and serviceable. Own ultrasound experimentations that feeds the image reconstruction team Design and execute experiment setups for transducer characterization (element sensitivity, bandwidth, cross-talk mapping, beam profile measurements) and ex vivo / phantom clinical testing. Acquire, process, and analyze RF and baseband signals for data quality assessment and benchmarking. Design simple boards and adapters as needed (monitoring, power/safety, interface/conditioning), and take them from prototype through a stable revision. Prototype quickly, then harden what works: wiring/harnessing, grounding, safety interlocks, and reliable integration across subsystems. Own practical test setups and documentation (fixtures, scripts, procedures) that make experiments repeatable and results comparable over time. What we’re looking for Strong hands-on EE background with experience building, debugging, and iterating on real systems in the lab. Solid understanding of signal processing fundamentals — knows what to measure, how to condition and digitize it, and how to evaluate signal quality in the context of an imaging system (SNR, bandwidth, dynamic range, artifacts). Comfortable spanning system integration + occasional design work (schematics/layout reviews or light PCB design) in a fast-moving environment. Ability to work at the boundary between hardware and algorithms: measure reality, communicate constraints, and help close gaps vs simulation. High agency and practicality: able to set up experiments, get trustworthy data, and unblock others on a lean team. Useful experience Analog/mixed-signal, or high-speed data capture experience; strong instincts for instrumentation and noise/debugging. Ultrasound or acoustic sensor handling: hydrophone calibration and field mapping, transducer impedance characterization, element-level sensitivity

GitAIGoRust
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set

PythonGitAIC++
F
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -87.7%

From $350K/yr

Quick readStrong listing-quality and freshness signals

About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Help Win New Business The Opportunity: We are scaling our dedicated Data Center practice, and we are looking for the person who will lead it. This is a founding commercial role. You will start as a team of one, owning the full sales motion end-to-end, and you will build the team around you as the practice grows. You will define how Flexport goes to market with hyperscalers, hardware OEMs, and the broader ecosystem. You will set the playbook, win the first marquee accounts, and hire the people who scale what you build. Reporting directly to the Regional General Manager, you will operate with a high degree of autonomy and direct access to executive leadership. This role features a 50/50 compensation model (Base + Uncapped bonus, with accelerators), with OTEs of $350k+, designed for strong leaders who take bets on themselves. Why This Role Is Different: The data center logistics market is not a standard freight problem. A single AI compute rack can cost more than $1 million and weigh up to 4,000 pounds. Racks contain Class 9 dangerous goods (lithium-ion batteries) and liquid cooling systems requiring specialized handling. Construction sequencing failures delayed 57% o

GitAIGoRust
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Account Security (AccSec) pod sits within the Accounts team (part of the broader Safety organization) and is at the heart of building a safe and civil platform for users of all ages. Account Security is especially challenging at Roblox because it is open to all ages. As the EM of AccSec pod (7+ ICs) you will work closely with peers throughout Roblox and the creator community to reduce the impact of account takeover (i.e. compromise) for Roblox players, creators, and the broader community. This is a role that requires entrepreneurship. We are looking for the EM of the AccSec pod to, in collaboration with product partners, to formulate a strategy for how to reduce ATO, reduce the impact of any remaining ATO and build community trust in the security of their accounts. Examples of past and current efforts include: Cryptographically binding user secrets to hardware backed secrets. Detecting client side tampering by Browser extensions. Integration with Passkeys. Post-login modeling for ATO detection. Reducing incentives for bad actors by making assets more recoverable. While this is a security team that focuses heavily on product and infrastructure changes to improve security, we also c

AWSGitAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a

PythonAWSGitAI
I
📍 United States· Full-time· Remote
✓ Quality checkedCompany trend -90.5%

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Why this role is on the menu Instacart’s product portfolio is expanding rapidly into new international markets, physical hardware like Caper carts, and retailer-facing platforms like Storefront Pro all in addition to Instacart’s core business. These initiatives need a s

MT
📍 Boise, ID - Main Site, United States· Full-time
✓ Quality checkedCompany trend +1150%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Position Overview: As a Staff/Principal Equipment Engineer within PROCT ADT, you will provide technical leadership to improve equipment performance, capability, and global manufacturing outcomes. You will lead complex, multi-site initiatives to resolve equipment challenges, enable technology node readiness, and drive yield and defectivity improvements. This role requires deep expertise, strong ownership, and the ability to influence across teams and suppliers to deliver measurable impact in cost, performance, and scalability. Responsibilities Lead global ownership of equipment performance across toolsets, driving improvements in stability, availability, matching, and efficiency. Direct complex, multi-fab root cause analysis for equipment-driven yield, defectivity, and reliability issues using data-driven and physics-based approaches. Provide technical leadership in tool hardware and chamber behavior to expand process capability and improve performance limits. Partner with Technology Development, integration, and process teams to enable node readiness, volume ramp, and disciplined change control. Influence equipment supplier strategy by leading technical engagements, driving design improvements, and aligning roadmaps to business needs. Drive global standardization through development, validation, and deployment of Best Known Methods (BKMs) across sites. Lead initiatives to improve cost of ownership, reduce variation,

AIRecruitment
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

AWSAzureRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o

AWSRestAIRust
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -88%

About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work

AWSRestAIRust
🔔

Get new hardware operations engineer jobs in United States by email

Daily job updates · Unsubscribe anytime