About the Team OpenAI’s Hardware organization develops system and infrastructure solutions tailored to the demands of advanced AI workloads. We work across the full stack—from silicon to system integration—partnering closely with internal teams and external vendors to define and deliver next-generation AI infrastructure. Our team focuses on defining scalable, high-performance system architectures and reference designs that balance performance, cost, and operational efficiency across rapidly evolving technologies. About the Role We are seeking a 3P Architect to define and drive rack- and cluster-level reference designs in collaboration with external partners. This role is responsible for translating workload requirements and system-level goals into concrete architectures, aligning partners on critical design attributes, and ensuring vendor roadmaps meet our infrastructure needs. You will work closely with performance modeling and internal architecture teams to evaluate tradeoffs, while owning the end-to-end definition and execution of third-party system designs. This includes identifying gaps in current technologies, driving vendor development, and shaping future infrastructure capabilities. This role requires strong system intuition, cross-functional leadership, and the ability to operate effectively across internal teams and external ecosystems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Define rack- and cluster-level reference architectures for AI infrastructure deployments. Translate workload requirements into clear system design specifications and partner deliverables. Collaborate with performance modeling teams to evaluate architectural tradeoffs and system behaviors. Align internal stakeholders and external partners on critical system attributes (performance, cost, power, reliability, scalability). Identify gaps in current technology offerings and dr
Jobs in United States
Partner Strategy And Operations Lead in San Francisco
1,034 active opportunities · Updated October 2026
Showing
15 jobs
Explore current partner strategy and operations lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. A majority of our users interact with our products in languages other than English, and our products must work seamlessly across languages, regions, and cultures. The Internationalization team builds the infrastructure that enables OpenAI products to ship globally by default. We develop the systems that power localization, international product launches, and high-quality global user experiences across all OpenAI products. About the Role As a Senior Software Engineer on the Internationalization team, you will build the systems that power localization and international product launches at OpenAI. You’ll work on the platform that manages product content, translation workflows, and localization infrastructure across our products. This role sits at the intersection of AI systems, developer platforms, and product infrastructure. In this role, you will Build and scale OpenAI’s localization, content, and experimentation platform used across OpenAI product teams, including open-source components: Develop AI-powered translation pipelines combined with human-in-the-loop review workflows. Design systems that reliably deliver localized product content across web and mobile apps. Build tools that enable linguists and localization teams to review and improve translations. Develop developer tooling that simplifies localization and internationalization workflows. Build and maintain internationalization libraries used across OpenAI products: Design systems that correctly handle numbers, currencies, dates, and pluralization across locales. Improve support for multilingual interfaces and right-to-left languages. Partner with product teams to improve the international readiness of new features. You might thrive in this role if you Have strong software engineering experience building backend or full-stack systems. Have familiarity with Java, React, MySQL, and cloud infrastructure p
About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for engineers to build the infrastructure that powers Codex agents in production. This role focuses on the systems that let models safely execute code, interact with tools, complete long-running tasks, and operate reliably and efficiently at scale. You’ll design and operate the infrastructure behind sandboxed execution, orchestration, stateful workflows, app-server and SDK boundaries, and model rollouts. You’ll work at the intersection of distributed systems, developer tooling, and AI, building primitives that make Codex faster, safer, more reliable, and easier for the rest of the organization to build on. What You’ll Do Design and build execution environments for AI agents, including sandboxing, isolation, and reproducibility. Develop systems for agent orchestration across multi-step, tool-using workflows. Build infrastructure for running, testing, and debugging code generated by models. Create state and memory systems that allow agents to persist context across long-running tasks. Optimize tokens, latency, reliability, and cost across Codex’s production fleet. Support model rollouts, capacity planning, and the core tradeoffs between quality, speed, and economics to manage a fleet of frontier agents at scale. Build shared platform capabilities that unblock product teams, partner teams, and open source Codex. Yo
About the Role We’re looking for a Procurement Enablement Lead to improve how employees and stakeholders navigate procurement at OpenAI. This role partners across Procurement, Finance, Legal, Security, Privacy, and Enterprise Technology to simplify workflows, improve guidance, and support scalable procurement experiences across the procurement lifecycle. You’ll translate procurement policies and operational requirements into clearer processes, better-enabled systems, and more intuitive employee experiences that reduce friction while strengthening consistency and controls as OpenAI continues to grow. This role is ideal for someone who combines operational judgment, process design, and strong cross-functional partnership skills. You should be comfortable working in evolving environments where systems and workflows are still being built and continuously improved. A key part of the role will be identifying opportunities to use AI and automation to streamline workflows, reduce manual work, and improve service delivery across Procurement operations. This role is based in San Francisco, CA. We use a hybrid work model of 3 days per week in the office and offer relocation assistance to new employees. In this role, you will: Partner across Procurement, Finance, Legal, Security, Privacy, and Enterprise Technology to improve how procurement work gets requested, routed, approved, and supported across the spend lifecycle. Help design and improve procurement intake, guidance, and workflow experiences that make it easier for employees and stakeholders to navigate procurement processes. Translate procurement policies, operational needs, stakeholder feedback, and operational insights into clear business requirements for workflows, automation, reporting, analytics, and process improvements. Partner with Enterprise Technology and tool owners to support workflow configuration improvements across procurement systems, including approvals, routing, SLAs, escalation paths, exception handlin
About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C
About the team OpenAI’s Education team is building products and experiences that help learners, educators, and institutions benefit from AI in ways that are rigorous, useful, and grounded in real learning outcomes. The work spans both consumer and B2B education, with close collaboration across engineering, learning science, design, data, and research. This team sits in a highly strategic investment area for OpenAI, with strong opportunities to shape how product ideas flow across consumer and institution-facing experiences. Some of our recent work: New Education Plugins for ChatGPT Work and Codex New tools for understanding AI and learning outcomes Education for countries Advancements in higher education Early product work - Introducing Study Mode About the role We’re looking for a product-minded Full Stack Engineer to help build OpenAI’s education products from the ground up. You’ll own end-to-end development across the stack, from early concepting and prototyping through production launch and iteration. This is an opportunity to work on a highly strategic, early-stage product area where engineering judgment, product sense, and customer empathy all matter. You’ll partner closely with leaders across the education org, including learning scientists, researchers, designers, and cross-functional partners, to turn emerging ideas into durable product experiences for schools, universities, and other education stakeholders. In this role, you will: Build and ship product experiences across the full stack for OpenAI’s education offerings Own projects end-to-end, from ideation and technical design through implementation, launch, and iteration Work closely with learning scientists and researchers to translate learning goals and evidence into product decisions Collaborate with design, data, and cross-functional partners to build thoughtful, high-quality user experiences Help define the engineering foundation for a growing education pod, including patterns, systems, and technical
About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work
About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal a dependable, scalable foundation for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between physical hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Security Engineer, Host Assurance to help build the trust foundations for bare-metal platforms across OpenAI’s global infrastructure. This is a deeply hands-on engineering role for a builder who can design, implement, and operate the core security infrastructure that establishes trust in hardware platforms before they are eligible to run workloads. Success in this role requires strong technical judgment, the ability to work comfortably at low levels of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—including novel hardware platforms and emerging deployment models– to make the secure path the easiest path. This role is well suited for engineers who enjo
About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o
About the Role The AI Deployment Manager (ADM) - Pilots is a customer-facing role responsible for leading structured, time-bound enterprise AI pilots from initial scoping through final executive readout. This role is focused on helping customers evaluate OpenAI’s products in real-world contexts, identify high-value use cases, and generate clear, decision-ready signals tied to business value. You will design and lead pilot engagements that drive activation, sustained usage, and measurable impact across ChatGPT Enterprise, Codex, and adjacent workflows. This includes partnering with customer stakeholders to define success criteria, guiding users from experimentation to real adoption, and translating pilot outcomes into clear recommendations that support expansion or purchase decisions. This role requires strong judgment, the ability to operate in ambiguity, and a consistent focus on connecting technical capabilities to business outcomes. You will regularly engage both executive stakeholders and working teams, adapting your approach to meet customers where they are and move them forward. In this role, you will: Own the design and execution of enterprise AI pilots, including scoping, cohort definition, and success criteria aligned to a clear commercial decision. Identify and prioritize a small set of high-impact use cases that can generate credible signal within a 30–45 day pilot. Drive activation and sustained engagement across pilot cohorts through targeted enablement, office hours, and workflow-level coaching. Monitor pilot performance and adapt in real time, diagnosing gaps in engagement, use case traction, or stakeholder alignment. Translate pilot signals into clear, executive-ready recommendations, including whether and how the customer should expand. Navigate customer constraints such as security, data access, and competing tools while maintaining pilot momentum. Partner closely with ADs, SEs, and customer stakeholders to align on scope, risks, and next steps. Ca
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking an experienced SoC Architect to lead the definition and development of next-generation custom AI silicon for edge deployments. This role will be responsible for shaping the architecture of highly efficient, high-performance SoCs optimized for machine learning inference and on-device intelligence. You will work cross-functionally with internal engineering teams and external ecosystem partners to translate product requirements into scalable silicon solutions, driving execution from concept through delivery. In this role you will: Define the architecture and technical roadmap for custom SoCs targeted for edge applications. Drive system-level tradeoff analysis across compute, memory, interconnect, power, thermal, and cost constraints. Architect energy-efficient ML compute subsystems optimized for inference workloads and real-world deployment environments. Collaborate with internal hardware, software, systems, and product teams to align architecture with platform needs. Partner with external silicon vendors, IP providers, and manufacturing partners to execute development plans. Lead hardware/software co-design efforts to maximize performance per watt and end-to-end system efficiency. Guide implementation teams through microarchitecture, RTL development, validation, and bring-up phases. Operate effectively in agile development environments and help teams deliver against aggressive schedules and milestones. You might thrive in this role if: Proven exper
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role The stakes of securing OpenAI increases as our internal coding and research becomes increasingly driven by autonomous AI agents. Compromising these agents could allow a cyber threat actor to compromise many other parts of the company. In this role, you would lead Preparedness work defending the security of our internal AI agents against insiders, Advanced Persistent Threats (APTs), or powerful AI agents. We’re looking for a strong hands-on technical executor with experience working directly with advanced cyber threat actors. In this role, you will: Develop and maintain threat models via which advanced attackers could compromise our coding assistants and automated security systems. Identify security investments that are especially critical to make in advance; for example, prioritizing by implementation lead-times, costs, and benefit. Partner with Security, Infrastructure, Research, Legal, and Preparedness to align on implementation plans and tradeoffs. Lead technical execution directly when needed, including prototyping controls, writing and reviewing software, and coordinating engineers across teams. Work with penetration testers to close gaps in defenses. You might thrive in this role if you: Are an exceptional hands-on technical executor. Have worked with advanced
About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a Software Engineer focused on building and scaling retrieval systems. You’ll work with a team of researchers and engineers to develop infrastructure that enables models to retrieve and act on the right information at the right time. This includes designing and operating indexing systems, retrieval pipelines, and serving layers. This work supports retrieval across OpenAI products and research, with direct impact on system performance, reliability, and scale. Responsibilities Build and scale retrieval infrastructure across indexing, serving, and query execution. Develop low-latency, high-throughput systems for real-time model interaction. Partner with research to productionize embedding and retrieval techniques. Support dense, sparse, and hybrid retrieval pipelines. Own system performance, reliability, and observability at scale. Collaborate across Pretraining, Inference, and Product teams to integrate retrieval end-to-e
Other cities to consider
More places hiring for this role
Get new partner strategy and operations lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime