Jobiba hiring network

Hardware Systems Planning Lead Jobs

701 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current hardware systems planning lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Peloton
📍 Taichung CityFull-time
29 days ago

ABOUT THE ROLE Peloton Electrical Engineers are responsible for the architecture, design, and testing of hardware systems for Peloton products. This early-career position requires working collaboratively with engineers and engineering managers to support product improvement and development initiatives across the entire product lifecycle, from new product introduction through sustaining engineering. As the team is highly dynamic with an exciting product roadmap, you will play a key role in ensuring electrical design robustness, manufacturability, time to market, and cost goals are all met. You will work multi-functional with many teams including Program Management, Product Management, Operations, Quality, Mechanical Engineering, Firmware Engineering, and Industrial Design. YOUR DAILY IMPACT AT PELOTON Under the supervision of senior engineers and project leaders, develop and support product subsystems, adhering to department and company standards and processes Observe and execute on engineering requirements rigorously Write and execute clear test plans at the design and production phases Troubleshoot and solve performance problems during development and in production Efficiently write and archive test reports Clearly document and communicate your work to senior engineers and technical leads Use data to drive engineering decision making Create and manage engineering changes for updates and improvements to existing products Collaborate with adjacent development teams to ensure design success, including compliance, firmware, and mechanical Work with drafting team to ensure appropriate product definition and design controls Communicate with outside vendors and partners effectively Travel throughout Asia and to/from the USA is expected to be 10%-15% YOU BRING TO PELOTON Bachelor’s degree in Electrical, Electronics, or Systems Engineering (Master’s degree preferred) Previous internship or similar work experience as an engineer in a design/develo

redisawsrest
View job →
O
OpenAI
📍 San FranciscoFull-time
27 days ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Product Manufacturing Engineer to drive manufacturing strategy and execution for next-generation AI hardware within the Silicon & Systems team, with a particular focus on PCBA manufacturing, assembly, process development, and production readiness. You will work closely with design engineering, systems engineering, operations, TPMs, contract manufacturers, suppliers, and other external partners to ensure new hardware products successfully transition from concept and prototype builds through NPI and high-volume production. You’ll own critical manufacturing initiatives, identify and resolve production risks, and help establish the processes and controls required to deliver complex hardware at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Drive manufacturing and quality initiatives to ensure product success from concept and early development through NPI, production launch, and scale. Lead manufacturing process development for next-generation AI hardware systems, partnering closely with design engineering, systems engineering, operations, TPMs, suppliers, and manufacturing partners. Own PCBA manufacturing development and production readiness, including assembly processes, process validation, manufacturing test, rework, yield improvement, and successful transition into volume production. Establish NPI manufact

awsrestai
View job →
O
OpenAI
📍 SingaporeFull-time
1mo ago

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Manufacturing Test Engineer to own and drive manufacturing test strategy, development, and execution for complex AI hardware systems. This role will define and implement test coverage across the product lifecycle, including ICT, functional circuit test, tray-level functional test, and system manufacturing test. You will work closely with hardware design engineering, diagnostic/software teams, manufacturing engineering, quality, and external system integrators and suppliers to translate product requirements into robust, scalable, and production-ready test solutions. You will also play a key role in reviewing test data, debugging failures, improving yield, and ensuring manufacturing test readiness from early development through volume production. In This Role, You Will: Define and drive the manufacturing test strategy for boards, trays, and system-level hardware assemblies across EVT, DVT, PVT, and production ramp. Develop and manage test coverage for: ICT / structural test FCT / board-level functional test Tray-level functional and integration test System-level manufacturing and bring-up test Partner closely with electrical engineering, system engineering, and diagnostic/software teams to define test requirements, review manufacturing test scripts and diagnostics, validate failure isolation needs, debug hooks, logging, and production screening strategies. Translate engineering requirements into practical, scalable manufacturing test plans that

pythonawslinux
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI is developing custom silicon to power the next generation of frontier AI models. We’re looking for experienced Design Verification (DV) Engineers to ensure functional correctness and robust design for our cutting-edge ML accelerators. You will play a key role in verifying complex hardware systems—ranging from individual IP blocks to subsystems and full SoC—working closely with architecture, RTL, software, and systems teams to deliver reliable silicon at scale. In this role you will: Own the verification of one or more of: custom IP blocks, subsystems (compute, interconnect, memory, etc.), or full-chip SoC-level functionality. Define verification plans based on architecture and microarchitecture specs. Develop constrained-random, directed, and system-level testbenches using SystemVerilog/UVM or equivalent methodologies. Build and maintain stimulus generators, checkers, monitors, and scoreboards to ensure high coverage and correctness. Drive bug triage, root cause analysis, and work closely with design teams on resolution. Contribute to regression infrastructure, coverage analysis, and closure for both block- and top-level environments. You might thrive in this role if you have: BS/MS in EE/CE/CS or equivalent with 3+ years of experience in hardware verification. Proven success verifying complex IP or SoC designs in industry-standard flows Proficient in SystemVerilog, UVM, and common simulation and debug tools (e.g., VCS, Questa, Verdi). Strong knowledge

awsrestai
View job →
O
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 United StatesFull-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

awsrestai
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t

awsrestai
View job →
A
Anyscale
📍 RemoteFull-time$171K – $211K/yr
20 days ago

At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale runs on a small, high-leverage IT function, and we're looking for an IT Engineer to own it. This is a hands-on engineering role, not a help desk or tier-1 support role. Roughly half the job is software: you will own and extend a set of internal automation and notification services wired into our directory and hardware systems. You should be comfortable owning a codebase and a GitHub repository, not only vendor admin consoles. The other half is the identity, endpoint, hardware, and vendor backbone that keeps a company of roughly 200 people, close to half of them engineers, productive across three public clouds. Reporting to the Head of Security and IT, you will work closely with security, engineering, and the people team. This role is based in the San Francisco Bay Area on a hybrid schedule. What You'll Do Own identity and access: Okta and Google Workspace administration, group-based authorization, 1Password Enterprise, passkeys, and automation driven provisioning and deprovisioning. Own and extend our internal automation services, and run them on scoped service accounts with sound security and hygiene. Own the endpoint fleet through mobile device management, including device trust posture checks, endpoint

azuregitmachine learning
View job →
N
Nuro
📍 Mountain View, California (HQ)Full-time
29 days ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We are seeking an experienced NPI Engineer to join our Autonomous Vehicle Hardware team. In this role, you will lead new product introduction activities for autonomous vehicle hardware systems, supporting the transition from prototype through pilot builds and into production. You will work cross-functionally with design engineering, manufacturing, quality, supply chain, and partners to ensure products are released with robust processes, scalable manufacturing plans, and clear readiness for launch. You will play a critical role in driving manufacturing readiness, launch execution, and cross-functional alignment for complex autonomous vehicle hardware programs. About the Work &

O
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a systems-minded engineer to help advance our kernel development, performance engineering, and hardware-software co-design capabilities, with a particular focus on AI-assisted workflows and tooling. This person will work at the intersection of kernel optimization, developer tooling, observability, and research infrastructure, helping us improve both how production kernels are built and optimized, and how future hardware-software systems are designed and evaluated. The role is ideal for someone who is excited by low-level performance work, but also sees AI and automation as powerful tools for accelerating engineering velocity. You will help define the future of kernel engineering in the era of AI-assisted development. In this role, you may: Build developer tooling and workflows that make kernel development and performance optimization faster, more scalable, and easier to debug, integrate, and deploy. Develop observability, diagnostics, and validation infrastructure that makes AI-assisted optimization systems more interpretable, reliable, and effective. Optimize production kernels end to end by formulating optimization problems, running search loops, analyzing bottlenecks, debugging generated implementations, and landing improvements into production. Design abstractions, interfaces, and automation systems that accelerate kernel optimization, correctness validation, and hardware-software co-design. Improve AI-assisted optimization systems for sp

awsrestai
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali

awsrestai
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

pythonawskubernetes
View job →
O
20 days ago

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo

awslinuxrest
View job →
🔔

Get new hardware systems planning lead jobs by email

Daily job updates · Unsubscribe anytime