About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
Jobs in United States
Hardware Operations Engineer in San Francisco
184 active opportunities · Updated October 2026
Showing
15 jobs
Explore current hardware operations engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
$226K – $285K/yr
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a Supply Chain Program Manager, you will own material readiness and supply chain execution for critical hardware programs spanning custom silicon, systems, memory, storage, networking, and rack infrastructure. You will work cross-functionally with Engineering, Strategic Sourcing, Manufacturing Operations, Finance, Planning, Quality, and external suppliers to develop and execute scalable supply strategies that support aggressive product development and deployment timelines. This role requires deep understanding of hardware supply chains, material planning, NPI execution, supplier management, and operational scaling in constrained and rapidly evolving environments. In this role you will: Material Readiness & Supply Planning - Own end-to-end material readiness across NPI and production phases, including building the necessary framework and processes for enablement. Drive supply planning and execution for long lead-time and constrained commodities including ASICs, HBM, DDR, SSDs, networking, optics, power, thermal, and mechanicals. Build and manage material readiness plans aligned to proto/pre-EVT, EVT, DVT, PVT, and mass production schedules. Monitor supply health, lead times, inventory positions, allocation risk, and capacity constraints. Drive shortage management, allocation mitigation, and recovery planning. Coordinate supply commits, forecast alignment, and supply continuity planning with suppliers and manufacturing partners. Cross-Functional Program Ma
About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req
About the Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We are seeking an experienced commercial attorney to serve as the primary legal partner supporting Robotics. In this role, you will provide practical, business-oriented legal counsel across hardware development, manufacturing, supply chain, procurement, and strategic commercial initiatives. You will work closely with Robotics leadership and cross-functional partners to help build scalable legal frameworks that enable innovation while thoughtfully managing risk. This is a unique opportunity to help shape the legal foundation of a rapidly growing robotics organization developing cutting-edge technologies. You will negotiate high-impact commercial agreements, advise on complex operational matters, and partner closely with technical and business teams to support the development and commercialization of next-generation robotics systems. This role is based in San Francisco, CA and requires in-person presence 4 days a week. In this role, you will: Serve as the primary legal partner supporting Robotics leadership and cross-functional teams. Provide practical, business-oriented legal advice to teams across hardware engineering, manufacturing operations, supply chain, procurement, finance, and operations. Draft, review, and negotiate complex commercial agreements, including supplier, manufacturing, development, consulting, licensing, procurement, and strategic partnership agreements. Advise on legal issues arising throughout the hardware development lifecycle, including manufacturing, supply chain operations, vendor relatio
About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, cloud services, mechanical engineering, electrical engineering, and product design to deliver reliable, production-ready devices at scale. Within Consumer Devices, Hardware Engineering eXperience, or HEX, is a new bootstrapped team building the environments, applications, compute, product-data systems, and workflows that let hardware engineers do their work without needing to troubleshoot the machinery underneath. HEX owns virtual engineering environments, HPC/GPU compute, storage, networking, licensing, MCAD/ECAD/CAE applications, PLM, product data, automation, validation, and support as one connected system. About the Role As a Staff PLM & Engineering Applications Engineer, you will be one of the first technical builders of HEX and the primary counterpart to the HEX lead. You will own the engineering-application and product-data side of the hardware engineering experience, with an initial focus on NX, Teamcenter, licensing, parts import, integrations, packaging, validation, and user workflows. This is not a traditional Teamcenter administration role and not a Corporate IT application-support role. You will take complex, fragile workflows and turn them into reliable engineering systems. This role is highly hands-on and systems-oriented. You will not inherit a mature environment and support queue. You will help build a fresh one, replacing manual setup guides, tribal knowledge, repeated support issues, and team handoffs with tested automation and reliable workflows. In This Role, You Will Own the technical architecture, deployment, configuration, integration, validation, and long-term operation of NX and Teamcenter. Build reliable workflows for parts import, product-data migration, metadata quality, BOMs, revisions, lifecycle states, and releas
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team The Connectivity Software Engineering team is responsible for enabling seamless, secure, and high-performance wireless connectivity across OpenAI’s products. We design and optimize Bluetooth, BLE, Wi-Fi, and emerging wireless technologies to ensure robust device pairing, network performance, and interoperability. Our work spans kernel drivers, system services, and user-level tools, with a focus on real-world performance, scalability, and reliability. About the Role OpenAI is seeking a Connectivity Software Engineer to design, implement, and optimize wireless connectivity features across our product ecosystem. You’ll work at the intersection of systems software, wireless standards, and hardware integration—building robust pairing and provisioning flows, debugging low-level protocols, and driving performance under real-world RF constraints. You will also support certification, field interoperability, and fleet-scale connectivity infrastructure. This role is based in San Francisco, CA . We use a hybrid work model of 4 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement, and debug Bluetooth/BLE and Wi-Fi features across kernel drivers, BlueZ/wpa_supplicant/hostapd, and systemd/D-Bus services Deliver robust pairing, bonding, and provisioning flows (GATT/GAP, LE Audio/LC3, WPA3/802.1X, captive portals, NAN) Optimize link performance: throughput, latency, jitter, roaming, coexistence (BT↔Wi-Fi), and power modes (TWT, WoWLAN) Build reliable network management using NetworkManager/nmcli, nl80211/cfg80211/mac80211, DNS/DHCP/mDNS, P2P/SoftAP Instrument and analyze with packet captures and tooling (btmon/hcidump, Wireshark, iperf, eBPF/perf, spectrum sniffers) Drive interoperability and certification readiness (Bluetooth SIG, Wi-Fi Alliance) and resolve field issues with root-cause fixes Contribute to OTA-safe configuration, telemetry, and diagnostics for fleet-scale operation You might thrive in
About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across silicon, embedded systems, operating systems, and cloud services to build reliable consumer devices and the novel platforms required to support them. We partner closely with research to bring advanced AI capabilities into the physical world. About the Role As an Operating Systems Engineer focused on on-device inference, you will design, develop, and ship the OS stack that makes advanced AI capabilities reliable, responsive, and energy efficient on consumer devices. Your work will span OS services and frameworks, inference runtime integration, model fitting, scheduling, and performance and power management. You’ll partner with research to adapt models to device constraints, make design decisions across the stack, and carry solutions from early exploration through integration and production. In this role, you will: Build the inference platform: Design and implement maintainable OS services, frameworks, and clear interfaces for inference execution, model loading and lifecycle, and resource management. Fit models to device constraints: Partner with researchers on quantization, runtime integration, and memory optimization to meet memory, compute, and energy budgets while evaluating model quality and product behavior. Coordinate system resources: Develop scheduling and resource policies that balance inference with other device act
Other cities to consider
More places hiring for this role
Get new hardware operations engineer jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime