About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o
Jobs in United States
Hardware Systems Planning Lead in United States
179 active opportunities · Updated September 2026
Showing
14 jobs
Explore current hardware systems planning lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The HR Business Partner (HRBP) team at OpenAI helps shape how our organization operates and performs. We work alongside senior leaders and their teams to design effective organizations, strengthen leadership capabilities, and help people do their best work. Our expertise spans leadership coaching, organizational design, talent strategy, employee relations, and change management. Our work is grounded in a deep understanding of OpenAI’s Contributions & Impact (C&I) culture and the needs of teams working at the frontier of AI and hardware. We operate with urgency, empathy, and sound judgment. We value sincerity over polish, collaboration over ego, and a relentless focus on meaningful impact. About the Role We’re looking for an experienced HR Business Partner to support OpenAI’s Consumer Device's team. This is an opportunity to partner closely with leaders and teams doing highly ambitious, multidisciplinary work. You’ll serve as a trusted advisor as the organization builds, collaborates, and evolves—helping create the conditions for people and teams to perform at their best. The role combines strategic advising with hands-on partnership. You’ll help leaders think through organizational questions, coach managers through complex situations, strengthen people practices, and make thoughtful trade-offs that balance the needs of our people, teams, and mission. This is an individual contributor role. Come build with us. Your Key Responsibilities Shape high-performing, resilient teams: Partner with leaders on organizational design, talent planning, and ways of working that support strong performance and sustainable teams. Advise with context and judgment: Develop a deep understanding of the organization and provide thoughtful guidance on people, leadership, and organizational matters. Collaborate across OpenAI: Work closely with fellow HRBPs, People centers of excellence, and cross-functional partners to create a cohesive employee experience rooted in hum
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Engineering Program Manager, you will help turn complex infrastructure strategy into executable programs across electrical, mechanical, controls, network, hardware, construction, commissioning, deployment, and operations workstreams. You will partner with research, hardware engineering, data center engineering, site development, supply chain, security, EHS, finance, legal, operations, and external delivery partners to bring OpenAI's infrastructure vision to life. About the Role We are looking for an Engineering Program Manager (EPM) to lead assigned infrastructure programs focused on production and non-production network integration, controls coordination, and the design and deployment of data hall or whitespace facilities. The EPM will support functional Directly Responsible Individuals (DRIs) across network, controls, structural, electrical, and mechanical disciplines. Key responsibilities include coordinating assigned workstreams and program controls, maintaining risks and interfaces, and supporting readiness within the network and data hall deployment track. The ideal candidate thrives on bringing structure to complex environments characterized by ambiguous technical requirements, large partner ecosystems, tight deadlines, and high operational stakes. This individual must be adept at keeping teams aligned on decisions, risks, dependencies, schedules, and readiness criteria, and escalating gaps or decision points when needed. Candidates should have a proven track record of managing technically challenging engineering programs across major lifecycle phases, including design, validation, procurement, construction, c
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Our Majors Hi-Tech team is expanding and we are seeking an elite Account Executive who doesn't just sell software but can understand the fundamental shift in how the world’s most sophisticated hardware and chip designers leverage data. In this role, you will lead the strategic expansion of Snowflake’s footprint within our most prestigious semiconductor accounts. You will navigate the intricate R&D, supply chain, and manufacturing lifecycles of the world's leading chipmakers and articulate why Snowflake is the essential foundation for their future. AS AN ACCOUNT EXECUTIVE AT SNOWFLAKE YOU WILL: Become an expert on Snowflake’s product and conduct discovery calls, customized demos, and presentations to prospective customers Build a deep familiarity with the Semiconductor/High-Tech industry and the specific workflows of the tech industry - understanding the pressures of global supply chains and complex R&D cycles. Land, adopt, expand, and deepen sales opportunities with accounts in your region Build executive alignment/relationships & develop champions within your accounts. Deeply embed yourself into the customer’s IT and R&D roadmaps, ensuring Snowflake is the backbone of their long-term innovation strategy. Collaborate & partner with Sales Engineering, Pro
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a Delivery Director, Capacity programs for our on-premises data center builds and neo cloud (GPU cloud) delivery programs. This is a high-visibility, execution-critical role sitting at the intersection of infrastructure engineering, capacity planning, vendor/partner management, and customer delivery. You will own the end-to-end delivery lifecycle for large-scale compute infrastructure — from initial site/capacity commitments through power, networking, and hardware bring-up, to production-ready GPU/compute capacity landing in the hands of internal teams or customers. You'll be the person who turns ambitious infrastructure roadmaps into predictable, on-time, delivery. RESPONSIBILITIES Own delivery of on-prem infrastructure builds — colocation expansions, power/cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing — coordinating across colo providers and partners, network engineering, hardware ops, and vendor teams. Drive neo cloud delivery programs — manage capacity delivery from GPU cloud and neo cloud partners (e.g., colocation/bare-metal/GPU cloud providers), including contract milestones, capacity ramps, SLAs, and go-live readiness. Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power/shell timelines, hardware lead times, logistics, and software/platform readiness into a single critical path.
From $126.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Data Center Engineer , you'll help us scale our Core/Edge Data Centers and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Technical Lead Data Center Engineer. You will: Develop and maintain the Core/Edge Data Center and hardware infrastructure to meet the large scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental life cycles. Own efforts to track and mitigate systemic issues preventing hosts from returning to service. Identify and solve critical problems and prevent them from re-occurring via root cause analysis and giving rec
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track
What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set
From $3.5M/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Account Security (AccSec) pod sits within the Accounts team (part of the broader Safety organization) and is at the heart of building a safe and civil platform for users of all ages. Account Security is especially challenging at Roblox because it is open to all ages. As the EM of AccSec pod (7+ ICs) you will work closely with peers throughout Roblox and the creator community to reduce the impact of account takeover (i.e. compromise) for Roblox players, creators, and the broader community. This is a role that requires entrepreneurship. We are looking for the EM of the AccSec pod to, in collaboration with product partners, to formulate a strategy for how to reduce ATO, reduce the impact of any remaining ATO and build community trust in the security of their accounts. Examples of past and current efforts include: Cryptographically binding user secrets to hardware backed secrets. Detecting client side tampering by Browser extensions. Integration with Passkeys. Post-login modeling for ATO detection. Reducing incentives for bad actors by making assets more recoverable. While this is a security team that focuses heavily on product and infrastructure changes to improve security, we also c
From $185K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Manager, Data Center Operations, you'll help us scale our Core Data Center and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Senior Manager of Data Center Operations. This will be a position based in Goodyear, AZ. You will: Develop and maintain the Core Data Center and hardware infrastructure to meet the large-scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental monitoring. Lead a growing team of data center engineers focusing on rack deployments, hardware troubleshooting and break-fix, and decommissioning. Identify and solve criti
About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We’re seeking a senior lawyer to participate in commercial legal strategy and execution across OpenAI’s fast-growing infrastructure portfolio. This is a cross-functional role that will partner closely with procurement, supply chain, partnerships, finance, and product teams to structure, negotiate, and manage the transactions that will support OpenAI’s long-term infrastructure ambitions. We’re looking for an experienced infrastructure transactions lawyer who thrives in ambiguity and wants to help define the commercial playbook for infrastructure efforts in the AI era. This role reports to the Associate General Counsel for infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. In this role, you will: Own commercial legal strategy and risk management for OpenAI infrastructure transactions. Draft, negotiate, and advise on complex agreements with infrastructure suppliers, manufacturers, distributors, and technology partners. Support strategic partnerships involving AI infrastructure and hardware supply chains. Develop frameworks for procurement, licensing, and collaboration across the infrastructure ecosystem. Partner with finance and operations teams to align contract terms with business and compliance requirements. Collaborate with policy and regulatory colleagues on issues impacting global supply chains, export controls, and manufacturing. Build scalable, efficient contracting process
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. In this role, you will: Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Collaborate with engineering and security teams to drive deployment of security enhancements and control changes across broad-scale infrastructure. Tackle high-impact projects such as checkpoint encryption, network isolation, secret management, and machine identity, while continuously raising the security bar for emerging AI workloads. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets to adapt to evolving challenges. You will thrive in this role if you have: Deep understanding of security principles, best practices, and common vulnerabilities. A proactive mindset, with the ability to identify and address secu
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
Other cities to consider
More places hiring for this role
Get new hardware systems planning lead jobs in United States by email
Daily job updates · Unsubscribe anytime