About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need
Jobs in United States
Hardware Systems Planning Lead in United States
179 active opportunities · Updated September 2026
Showing
15 jobs
Explore current hardware systems planning lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities Co-design future hardware for programmability and performance with our hardware vendors Assist hardware vendors in developing optimal kernels and add support for it in our compiler Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory h
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will help develop and validate next-generation server platforms that power a reliable, high-performing, and cost-efficient infrastructure at scale. You will work across platform bring-up, firmware qualification, hardware validation, fleet integration, and performance optimization to support large-scale production deployments. You Will: Bring-up & Sustaining: Drive key aspects of the hardware development lifecycle, including feasibility studies, hardware bring-up, validation, deployment, and ongoing production support. Platform Optimization: Perform platform integration, performance characterization, and system-level debugging across compute infrastructure, focusing on hardware optimization, driver tuning, and thermal/power efficiency. Hardware Validation: Develop and execute rigorous evaluation and stress-testing strategies for server platforms to ensure reliability and performance under production-scale workloads. Firmware & Fleet Enablement: Support BIOS/BMC firmware qualification, hardware health monitoring, and automation tooling for firmware deployment and lifecycle management. Vendor & Cross-Functi
From $79.5K/yr
Location Details: Santa Clara, CA At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is an in-office position and you’ll be expected to work full-time in an office location, and therefore must live within commuting distance from your assigned office. You will work from this office beginning on your first day. Join our team Join GoDaddy's Global IT Support team as a Desktop Support Technician in our Santa Clara office, where you'll be the face of IT for employees who rely on technology to do their best work every day. In this hands-on, 100% on-site role, you'll provide walk-up and escalated technical support, troubleshooting hardware, software, workplace technology, and connectivity issues while delivering an exceptional customer experience.You'll collaborate closely with a globally distributed team of support professionals across North America, EMEA, and APAC, helping resolve both everyday technical requests and high-priority incidents in a fast-paced environment. Success in this role is driven as much by your communication, empathy, and problem-solving skills as your technical knowledge, making it an excellent opportunity for someone who enjoys helping people and learning new technologies.If you're passionate about customer service, curious about how technology works, and looking to grow your career in enterprise IT, we'd love to hear from you. What you'll get to do... Provide front-line technical support to employees by diagnosing and resolving hardware, software, operating system, peripheral, and account-related issues, ensuring minimal disruption to productivity. Manage and prioritize incidents and service requests through the IT ticketing system, taking ownership from initial intake through troubleshooting, resolution, and follow-up communication. Configure, deploy, maintain,
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox is seeking a Hardware and Testing Lab Tecnician (Tech Support Specialist III) to support our automated testing device farm, a rapidly scaling on-premises hardware environment that enables automated testing across mobile, desktop, console, and specialty devices. This infrastructure is critical to Roblox's ability to ship quality software rapidly. This role sits within Corporate Engineering's Client Services team and is responsible for the physical and operational lifecycle of hundreds of test devices, from procurement and rack installation through daily health maintenance, break/fix support, and end-of-life recycling. You will work closely with Engineering to understand device configurations and testing requirements, while owning the hands-on execution that keeps these labs running reliably. Success in this role requires someone who takes full ownership of their work, holds themselves accountable, and can be counted on to follow through. This is a fully on-site role in our San Mateo, CA offices. You Will: Receive, asset-tag, inventory, and physically install devices (phones, tablets, Macs, PCs, and potentially consoles/VR) into IDF rack environments. Configure devices at t
About the Team OpenAI's Legal team plays a crucial role in furthering OpenAI's mission by tackling innovative, fundamental legal issues in AI. If you're passionate about doing significant and unique work as a technology lawyer, this team is for you. The team comprises legal professionals from diverse fields, including technology, privacy, IP, corporate, cybersecurity, employment, tax, regulatory, and litigation. About the Role We’re growing our world-class Legal team and seek an experienced counsel to lead our global hardware IP portfolio initiatives, including patent, trademark and other intellectual property matters related to our business. This role is highly cross-functional across OpenAI, including work across our Legal, Communications, Global Affairs, Product, Research and Executive teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own global hardware IP initiatives, including setting and executing on the strategic direction and management of our hardware IP portfolio. Create scalable processes internally and externally with outside counsel to build, maintain and protect our hardware IP portfolio. Developing and maintaining internal hardware IP policies and programs. Engaging externally on hardware IP policy issues. Advising on strategic hardware IP deals. Advising on hardware IP issues, ranging from patent, trademark, trade secret to open source. Developing and building internal AI expertise, processes and tools to facilitate legal team work. Experience advising on complex technology transactions and inbound technology licensing. You might thrive in this role if you: Have at least 10+ years of combined hardware IP experience at innovative technology companies and law firms. Have a JD and license or qualification to practice in CA. Have a strong sense of ownership, are inquisitive and enthusiastic about technology, enjoy being con
From $116K/yr
We're looking for someone to join the Datadog Procurement team and help expand the Strategic Sourcing group. Make an impact by being a trusted subject matter expert in several buying categories and continuing to prove the value that Strategic Sourcing brings to the organization. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic sourcing for enterprise SaaS and IT categories - owning net-new purchases and complex, high-spend renewals through a documented, collaborative evaluation process, with action plans built well ahead of renewal dates. Own software license and subscription strategy, right-sizing entitlements, minimizing shelfware, and replacing last-minute buying with forward-looking renewal and demand planning. Lead the IT hardware procurement team responsible for purchasing, sourcing, and replenishment (laptops, peripherals, networking, office AV), balancing demand with headcount and multi-region growth while leveraging OEM/distributor relationships and buying scale. Lead the transformation of IT Procurement (hardware and SaaS licenses) from transactional to strategic - standardizing processes, introducing supplier-management discipline, and establishing operational metrics and KPIs. Develop strong, trusted relationships with stakeholders and collaborate on strategizing the purchase plan, negotiating pricing and other business terms with vendors on net-new purchases and complex, high-spend renewals, partnering cross-functionally with Procurement Operations, Legal, Security, IT, and Office Operations to ensure compliant, scalable execution. Create pricing models based on vendor proposals to quantify various buying scenarios related to our future growth and current demand Manage and develop a direct report, with
About the Role OpenAI’s Applications organization spans rapidly growing Consumer, Enterprise, Developer, and Hardware businesses. As we scale across EMEA, APAC, and LATAM, we’re building the financial and operating model that will help our regional teams make faster, better decisions. We are hiring a senior Strategic Finance partner to help shape that work. You’ll serve as a trusted advisor to regional GMs, connecting market strategy with business performance and guiding where and how we invest. By bringing together insights across growth, revenue, OpEx, headcount, and marketing effectiveness, you’ll help leaders understand market health, set priorities, and drive measurable outcomes. You’ll create the frameworks, operating rhythms, and decision processes needed to scale our international business, while contributing to broader Strategic Finance priorities across Marketing. This role is based in our San Francisco HQ. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Periodic international travel is expected. In this role, you will: Own an integrated view of OpenAI’s international performance—spanning countries, regions, and Consumer, Enterprise/API, and Developer businesses—to guide growth, resource allocation, and investment decisions. Partner with regional GMs and functional leaders to shape country strategies spanning Product, Sales, Growth and Brand Marketing, Partnerships, Communications, and Policy. Lead regional financial planning and forecasting, tracking budgets, headcount, capacity, investments, and risks against the company plan. Build and operationalize investment frameworks across markets, channels, and programs, evaluating expected returns, opportunity costs, and tradeoffs to recommend what to fund and when. Create regional and country scorecards and lead regular business reviews that track key metrics, surface risks, drive decisions, and ensure follow-through. Design decision rights and escal
About the Team Like every team at OpenAI, the Marketing team contributes to our broader mission of ensuring responsible and widespread adoption of artificial intelligence With that aim in mind, we are responsible for developing and executing strategies that drive awareness, engagement, and usage for OpenAI’s products and platform amongst our core audiences. We take a data-driven approach to understand our customers' needs and challenges, ensuring that their voices are reflected in product development and messaging. We then partner closely with Product, Engineering, Research, Comms, and Design teams to create a cohesive customer experience across all our channels. Our focus extends beyond just promoting product features; we aim to provide valuable insights and resources that help our users make the most out of AI technologies. About the Role As a Product Marketing Manager for Partner Marketing on the Hardware team, you will help shape the partnership strategy for a new product ecosystem and bring that ecosystem to market with the right partners, programs, and commercial motions. This role calls for a strategic, relationship-oriented marketer who can identify high-potential partners, understand their priorities, and turn shared opportunities into durable, high-impact go-to-market plans. You will work closely with Product, Design, Engineering, Integrated Marketing, Sales, Comms, Finance, Legal, and partner-facing teams to define how first-, second-, and third-party partners participate in our hardware ecosystem. You will lead the marketing and commercial aspects of partnerships, support joint launches, and ensure partner needs are understood internally while OpenAI’s product, brand, and customer experience standards are represented externally. This role reports into the VP of Product Marketing, Hardware, and offers a unique opportunity to build the partner marketing foundation for a new category of AI-powered devices. In this role, you will: Help craft and implement th
Other cities to consider
More places hiring for this role
Get new hardware systems planning lead jobs in United States by email
Daily job updates · Unsubscribe anytime