About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with research, software, and external hardware partners to shape the next generation of AI systems, from silicon through full-scale deployments. Our team focuses on understanding and optimizing performance across the full system stack—ensuring that architectural decisions are grounded in rigorous, quantitative analysis of real-world workloads. About the Role We are seeking a Performance Modeling Lead to build and lead a small, high-impact team responsible for answering forward-looking architectural questions across AI infrastructure systems. You will develop modeling frameworks and methodologies to evaluate system-level tradeoffs and guide key design decisions. Your work will directly influence reference architectures, vendor designs, and long-term infrastructure strategy. This role sits at the intersection of AI workloads, system architecture, and quantitative modeling, and requires strong technical judgment, ownership, and the ability to translate complex analysis into clear, actionable guidance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Build and own a performance modeling framework/toolchain to evaluate AI systems across multiple levels of abstraction. Analyze and quantify architectural tradeoffs across compute, memory, networking, storage, and system topology. Develop performance models to guide decisions on: scale-up vs. scale-out architectures interconnect and network design memory hierarchy and system balance. Translate modeling outputs into clear recommendations for internal teams and external hardware vendors. Influence reference designs and vendor roadmaps through data-driven insights. Partner closely with machine learning, systems, and hardware teams to understand workload characte
Jobs in United States
Solution Engineering in United States
2,195 active opportunities · Updated October 2026
Showing
15 jobs
Explore current solution engineering jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate the software used to manufacture, qualify, and deliver our hardware from the factory. You will work across the stack to create the infrastructure that runs internally and externally to coordinate all aspects of the production process. You will create the critical tools and procedures to execute, capture, process, and present the data resulting from the end to end assembly and validation of our hardware across multiple vendors and sites. This role is hands-on and high-ownership. You will work closely across teams both internal and external to define the standards that will be used across our products to ensure the velocity and quality of our 1P hardware. You will own the implementation, deployment, and output of these systems as well their continued maintenance and SLAs. Location: San Francisco, CA (Hybrid: 3 days/week onsite). Relocation assistance available. In this role, you will: Design, develop, and maintain the software infrastructure for manufacturing process execution and data export. Own integration across internal customers and vendor systems and processes. Build and maintain the CI, release, and delivery pipeline of tooling to external partners. Build and maintain internal systems to ingest, process, deliver, and visualize critical data for internal teams and systems. Build system health monitoring, telemetry, remote d
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a highly skilled Physical Design Engineer with deep expertise in physical design and methodology. This individual contributor role sits within our physical design team and is central to delivering power, performance, and area (PPA) optimized datapath and interconnect solutions for next-generation AI accelerators. You’ll work closely with RTL designers to define and execute on physical design strategies. You will develop tools, flows and methodologies to increase team productivity. Your work will directly impact silicon’s performance and cost efficiency, as well as the team’s execution velocity and quality. In this role, you will: Develop, build and own tools, flows and methodologies for physical implementation Own physical implementation of floorplan blocks from floorplanning to final signoff Collaborate with RTL designers to drive optimal block implementation solutions Analyze and optimize design for timing, power, and area trade-offs, working in collaboration with EDA vendors and ASIC partners Qualifications: BS w/ 4+ or MS with 2+ years or PhD with 0-1 year(s) of relevant industry experience in physical design and methodology development Demonstrated success in taping out complex silicon designs Hands-on experience with block physical implementation and PPA convergence Strong coding experience with python, bazel, TCL Strong experience building physical design tools, flows and methodologies Strong understanding of microarchitecture, RTL design,
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will work on the systems software strategy and execution that brings new AI silicon from first power-on to a fully integrated system running production-representative models at expected functionality and performance. You will define how software exercises and validates compute, memory, interconnect, and I/O subsystems, then build the diagnostics, automation, and observability needed to find issues quickly. This role sits at the center of silicon, firmware, platform, systems, and workload teams. You will turn hardware specifications and performance targets into an end-to-end bringup plan, drive cross-functional debug, and establish the stress and regression infrastructure that makes each new platform reliable across operating environments. In this role, you will: Contribute to the end-to-end software bringup and validation strategy for new silicon and first-party systems. Define software-driven test coverage across compute, memory, interconnect, I/O, and their system-level interactions. Build diagnostics, test automation, telemetry, and regression infrastructure that accelerate first-silicon learning and issue isolation. Lead bringup from initial silicon arrival through board and system integration, docking, runtime enablement, and model execution. Design stress tests that characterize reliability, performance, and stability across workloads and operating conditions. Translate architecture specifications and performance models into measurable acceptance crit
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the low-level device runtime that turns compiled programs into efficient, functional and performant execution on OpenAI’s custom AI accelerator. This software will schedule kernel launches, manage device memory and address spaces, coordinate synchronization, and expose reliable abstractions to higher-level runtimes and frameworks. You will work at the boundary of software and hardware, partnering with compiler, kernel, architecture, verification, and silicon teams to define interfaces and validate behavior. You will also use and improve event-based, cycle-accurate simulation to develop runtime capabilities before silicon is available, diagnose performance and correctness issues, and guide hardware-software co-design. In this role, you will: Design and implement the low-level device runtime for OpenAI custom silicon. Build kernel-launch scheduling, command submission, queueing, dependency tracking, and completion handling. Manage device memory spaces, allocation, virtual-to-physical mappings, data movement, and lifetime across concurrent workloads. Implement synchronization primitives, events, barriers, streams, and ordering guarantees that are correct and efficient. Define clean interfaces between the runtime, drivers, firmware, compiler-generated code, kernels, and higher-level execution systems. Use event-based, cycle-accurate simulators to develop, validate, debug, and performance-tune runtime behavior before and after silicon availability. Di
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking an experienced ASIC Package Signal Integrity / Power Integrity Engineer to drive electrical architecture, modeling, optimization, and validation for the most advanced AI/HPC silicon and package design. This role focuses on high-speed SerDes and memory channel architecture, advanced 2.5D/3D package SI/PI, substrate to package co-design, power-delivery-network optimization, electromagnetic modeling, and simulation-to-measurement correlation. The ideal candidate has strong hands-on experience with high-speed channel and PDN analysis across ASIC packages, interposers, substrates, and power-delivery structures, and can translate simulation results into practical design requirements for interposers and package substrate design optimization. The engineer will work closely with ASIC, package, system, mechanical, thermal, power, and silicon validation teams from early architecture and feasibility studies through production bring-up. In this role you will Own SI/PI architecture and analysis for advanced AI ASIC packages from early feasibility studies through production. Develop and optimize high-speed electrical channels for 200G/400G SerDes, PCIe, HBM, DDR, and chiplet/die-to-die interfaces. Perform package, interposer and substrate modeling using 2D/3D electromagnetic solvers. Define and optimize package stack-ups, transmission-line structures, via transitions, breakout structures, return paths, ground shielding, bump maps, and ball maps based on SI/PI re
We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As an AI Accelerator Systems Software Technical Program manager at OpenAI, you will help bring our chips/system hardware roadmap to life, navigating an array of technical and partnership challenges. We’re looking for people excited to push the frontiers of computing by navigating technical explorations and are passionate about building. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage the end-to-end software development from design to implementation for our AI acceleration systems, working across technical, cross-functional and external stakeholders Lead planning and scheduling of AI system software designs with our strategic partners and vendors Coordinate and lead internal resources and communication for efficient interaction with partners and vendors. You might thrive in this role if you: Have experience as a software technical program manager for data center system products (server, GPU, TPU, networking, storage and so on) taking products from concept to volume in a data center environment ensuring the systems scale with high quality Know end-to-end software development program management techniques from concept, design, production, deployment into the data center Want to help design some of the world’s largest supercomputing systems, working at the edge of complex hardware challenges Enjoy working with and enabling world-clas
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for an experienced systems software engineer to help define and build the host software stack for our custom next-generation AI systems. You will work close to the hardware on performance-critical software, including Linux kernel drivers, high-throughput I/O paths, and system-scale networking and RDMA. This role spans architecture, implementation, platform bring-up, debugging, and performance optimization. You will work across hardware and software boundaries to make new systems usable end to end, from low-level device interfaces through userspace tooling and production validation. In this role you will: Design, implement, and debug host-side systems software for AI infrastructure, including Linux kernel drivers and supporting userspace components. Build and optimize software paths for high-throughput, low-latency communication, including RDMA and related networking functionality. Develop software around PCIe, DMA, NICs, accelerators, memory movement, and device interaction. Bring up new hardware platforms and diagnose complex issues across kernel, firmware, networking, and hardware boundaries. Build tooling for integration, testing, diagnostics, observability, qualification, and performance characterization. Collaborate with hardware, networking, and platform teams to define interfaces and integrate new capabilities. Work with external vendors where needed to integrate technologies and drive issues to resolution. Contribute across the systems sof
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Security Engineer to join our First-Party Hardware team. In this role, you will own the end-to-end security foundation for OpenAI's first-party AI hardware systems, working across hardware security, embedded security, system security, and practical deployment at data center scale. You will partner with silicon, hardware, firmware, infrastructure, manufacturing, operations, and security teams to define and deliver system-level device trust. This includes boot integrity, device identity, provisioning, attestation, management-plane security, storage encryption, debug controls, firmware update and recovery, RMA, and decommissioning. You will be accountable for turning threat models into requirements, requirements into implementation, and implementation into validation evidence that can support launch decisions. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Own security requirements, threat models, validation strategy, and launch-readiness evidence for first-party hardware platforms from early design through production deployment. Design and review secure boot, measured boot, roots of trust, platform firmware resilience, firmware signing, recovery, and anti-rollback strategies across heterogeneous devices. Own device identity, provisioning, enrollment, attestation, certificate lifecycle, and key-management requirements across manufacturing and data center bring-up. Harden management
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a System Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate low-level system software for the manageability and health of OpenAI's first-party AI hardware systems. You will work across BMC, Linux, firmware interfaces, automation infra, boot and recovery, hardware diagnostics, telemetry, host and platform drivers, network software interfaces, and manufacturing and fleet readiness. A major part of this role is owning the acceptance path for partner-delivered system software: defining requirements, reviewing code and artifacts, reproducing builds, building tests, pushing fixes, and producing the evidence needed for launch decisions. This role is hands-on and high-ownership. You will write and review low-level software, debug issues across hardware and software boundaries, build infra and automation to test and manage devices in lab, guide partner deliverables, build validation evidence, and help carry platforms from bring-up through production deployment. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Design, develop, and maintain low-level firmware and system software for first-party AI hardware manageability, including BMC software, Redfish services, gNMI telemetry, firmware update and recovery flows, BIOS/UEFI interactions, platform drivers, and hardware diagnostics. Own integration and acceptance of partner and ve
Other cities to consider
More places hiring for this role
Get new solution engineering jobs in United States by email
Daily job updates · Unsubscribe anytime