Jobs in United States

Infrastructure Team Manager in United States

1,503 active opportunities · Updated October 2026

Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Codex Research team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of the Codex Research team, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measu

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team IT Systems Operations serves as the operational control layer connecting Security, Engineering, and Employee Technology platforms. The team ensures employee-facing systems, identity workflows, and enterprise applications operate in a predictable, governed, and continuously reliable manner as the organization scales. Beyond implementation, the team establishes structured operating patterns that ensure platform changes, access models, integrations, and lifecycle workflows evolve safely and consistently across the enterprise environment. About the Role This role acts as an operational owner for identity-connected enterprise SaaS platforms and system change controls supporting OpenAI’s compliance requirements. The engineer will be responsible for ensuring that production configuration, access models, and platform changes remain compliant, auditable, and consistently operated through defined controls. You will own and operationalize controlled system and workflow changes across identity platforms, SaaS applications, collaboration tooling, and enterprise infrastructure. You will partner closely with Security, Platform Engineering, and IT Support Operations to: Ensure identity and access workflows behave consistently across systems Implement structured rollout and configuration practices for enterprise applications Improve visibility and traceability of system changes impacting employee workflows Translate operational requirements into durable automation and policy-aligned implementations Success in this role requires not only strong engineering capability, but also sound operational judgment, disciplined documentation practices, and effective cross-functional collaboration. In this role, you will: Enterprise SaaS & Identity Platform Ownership Own administration and operational stewardship of enterprise SaaS and identity-connected platforms, ensuring configuration integrity, access governance, and compliance with defined control requirements. Own onboard

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Learn more about OpenAI’s approach to safety. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. About the Role At OpenAI, we're dedicated to advancing artificial intelligence, and we know that creating a secure and reliable platform is vital to our mission. That's why we're seeking a software engineer to help us build out our trust and safety capabilities. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. Your Responsibilities: Architect, build, and maintain anti-abuse and content moderation infrastructure designed to protect us and end users from unwanted behavior. Work closely with our other engineers and researchers to utilize both industry standard and novel AI techniques to measure, monitor and improve AI models’ alignment to human values. . Diagnose and remediate active incidents on the platform and build new tooling and infrastructure that address the root causes of system failure. You might thrive in this role if: You have built and run production services in a high growth, rapidly scaling environment. You can debug live issues and restore systems quickly. You have worked on content safety, fraud, or abuse, or are motivated and excited to work on present-day (“now-term”) AI safety. You have experience with Python or with modern languages such as C++, Rust, or Go, and are able to quickly ramp up on Py

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We’re looking for a GPU Inference Engineer to contribute to improvements in model serving efficiency for our Robotics research. This is a high-impact role where you’ll drive initiatives to optimize inference performance and scalability. You’ll also be engaged in model design, to help assist our researchers in developing inference-friendly models. This role is critical to scaling the team’s broader goals - it will directly enable leadership to focus on higher-leverage initiatives by building a stronger technical foundation. In this role you will: Perform engineering efforts focused on improving model serving, inference performance, and system efficiency Drive optimizations from a kernel and data movement perspective to improve system throughput and reliability Partner closely with research and product teams to ensure our models perform effectively at scale Design, build, and improve critical serving infrastructure to support Robotics growth and reliability needs You might thrive in this role if you: Have deep expertise in model performance optimization, particularly at the inference layer Have a strong background in kernel-level systems, data movement, and low-level performance tuning Are excited about scaling high-performing AI systems that serve real-world, multimodal workloads Can navigate ambiguity, set technical direction, and drive complex initiatives to completion This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. About OpenAI OpenAI i

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer you will join our OpenAI engineers and researchers in building, operating and securing transformational AI technologies. This role will focus on all aspects of Detection & Response but with a strong emphasis on detecting insider threats and influencing controls to safeguard OpenAI's most sensitive assets. In this role, you will: In this role, you will: Innovate on Detection and Response infrastructure to engineer and automate end-to-end detection and investigation workflows. Develop, measure, and tune detection rules to ensure effective and sustainable operations. Drive projects across OpenAI’s technology stack with a focus on insider threats, ranging from access abuse and intellectual property theft to novel risks emerging within AI infrastructure. Partner closely with cross-functional stakeholders, including HR, Legal, and peer investigative teams, providing technical expertise and evidence to support investigations. Collaborate on cutting-edge AI research, and use AI to improve OpenAI’s Security posture. You might thrive in this role if you: 5+ years experience working in a detection/response or insider-risk role.. We are seeking mid-level and senior candidates. You have broad familiarity with operating systems and platforms such as macOS, Windows, Linux, and Kubernetes, along with experience in cloud infrastructure. Knowledge of modern adversary tactics and attack paths, data exfiltration techniques, and h

PythonAWSKubernetesLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the core infrastructure behind OpenAI’s monetization and ads systems. In this foundational role, you’ll architect and implement distributed systems that power OpenAI’s monetization stack—focusing on reliability, performance, privacy, and large-scale operation. You’ll work across backend, systems, and platform layers to define and implement 0→1 infrastructure, partnering closely with Product, Design, and Research to shape the future of monetized AI experiences. Your work will enable both internal and external teams to build on safe, scalable, and robust monetization primitives. This role is exclusively based across our San Francisco & Seattles sites. We offer relocation assistance to new employees. In this role, you will: Design and build the foundational backend and infrastructure powering OpenAI’s monetization and ads systems Architect large-scale distributed systems that

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Role The Engineering Acceleration team builds and operates the foundational systems that engineers use to build, test, and ship ChatGPT, the API, and OpenAI's infrastructure. We are looking for an engineer to help evolve OpenAI's build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, and software quality. You will work on the systems that determine how quickly and confidently engineers can move: Bazel-based builds, Buildkite pipelines, test selection, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to make OpenAI one of the most productive engineering organizations in the world while preserving a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping useful systems instead of fighting infrastructure. In This Role, You Will Own and evolve Bazel-based build and test workflows across a large, polyglot monorepo. Design and maintain Starlark rules, macros, toolchains, and integrations that make builds reproducible, hermetic, and easy for product teams to adopt. Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, test sharding, retry behavior, and flake isolation. Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling. Improve local development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly without learning every detail of the build stack. Operate and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, and remote cache/exec

TypeScriptPythonAWSDocker
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products, including next-generation ads experiences, that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers, advertisers, and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring new ad experiences into real-world systems across OpenAI surfaces at global scale, including thoughtfully integrating them into the core ChatGPT experience. About the Role We’re looking for an Android Engineer to help build the native Android experiences and client-side systems that power how ads are structured, rendered, and delivered across OpenAI’s ads ecosystem. This role sits within the Ads Formats team, which owns the creative rendering and presentation layer for next-generation ads experiences across different surfaces and media types. You’ll help build the Android infrastructure and tooling that support formats such as text, image, video, native, conversational, and interactive ads, while ensuring they render reliably and perform efficiently across platforms. You’ll work closely with backend, Product, Design, Research, and Safety partners to shape Android architecture that supports user-first, privacy-preserving monetization experiences within ChatGPT and OpenAI’s broader mobile ecosystem. In th

AWSRestAIKotlin
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the core monetization and ads systems at OpenAI. This is a foundational role responsible for designing and implementing the infrastructure, APIs, and user-facing experiences that will power OpenAI’s next-generation monetization products—including ads. You’ll work across the full technical stack to architect, build, and ship 0→1 systems that are robust, safe, and scalable. You will collaborate deeply with Product, Design, and Research to define the future of monetized AI experiences and ensure these systems meet OpenAI’s highest standards for safety, privacy, and policy alignment. This role is exclusively based across our San Francisco and Seattle sites. We offer relocation assistance to new employees. In this role, you will: Design, build, and scale the core infrastructure behind OpenAI’s monetization and ads products Develop advertiser-facing APIs and tools that enable the cre

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI is developing custom silicon to power the next generation of frontier AI models. We’re looking for experienced Design Verification (DV) Engineers to ensure functional correctness and robust design for our cutting-edge ML accelerators. You will play a key role in verifying complex hardware systems—ranging from individual IP blocks to subsystems and full SoC—working closely with architecture, RTL, software, and systems teams to deliver reliable silicon at scale. In this role you will: Own the verification of one or more of: custom IP blocks, subsystems (compute, interconnect, memory, etc.), or full-chip SoC-level functionality. Define verification plans based on architecture and microarchitecture specs. Develop constrained-random, directed, and system-level testbenches using SystemVerilog/UVM or equivalent methodologies. Build and maintain stimulus generators, checkers, monitors, and scoreboards to ensure high coverage and correctness. Drive bug triage, root cause analysis, and work closely with design teams on resolution. Contribute to regression infrastructure, coverage analysis, and closure for both block- and top-level environments. You might thrive in this role if you have: BS/MS in EE/CE/CS or equivalent with 3+ years of experience in hardware verification. Proven success verifying complex IP or SoC designs in industry-standard flows Proficient in SystemVerilog, UVM, and common simulation and debug tools (e.g., VCS, Questa, Verdi). Strong knowledge

AWSRestAIRust
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About The Team Our mission is to bring OpenAI products to life for every customer. Demo Experience equips customer-facing teams with the experiences, systems, and confidence to make frontier capabilities tangible, relevant, and trustworthy. OpenAI’s products and customer needs are evolving rapidly. Demo Experience closes the gap between a frontier capability and a credible customer experience—making new capabilities understandable, demonstrable, and reusable quickly at scale. Working across Product, Engineering, Marketing, Operations, and GTM, we turn recurring customer needs into reusable capabilities and raise the standard for every customer conversation. About The Role Demo Experience Engineers work at the intersection of product engineering, technical storytelling, and GTM execution. You will own ambiguous, high-leverage problems end to end—from building agentic prototypes to creating the infrastructure and self-service tools that make them reliable and reusable. Your work will help customer-facing teams move faster, reduce avoidable failures, and translate frontier product capabilities into clear customer value. You will also turn recurring patterns from customer-facing work into product feedback, launch-readiness improvements, and scalable systems. Success in this role means teams can demonstrate new capabilities sooner and with greater confidence. Recurring requests become reusable capabilities instead of one-off work. Demo experiences are accurate, reliable, and safe. Insights from customer-facing work improve product and readiness decisions. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own end-to-end demo readiness for new and priority product capabilities, including environments, integrations, synthetic data, evaluations, reliability checks, and fallback paths. Build compelling prototypes, LLM agents, and reference flows that mak

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. Analyze inference workloads end to end across applications, models, and fleet infrastructure. Enhance tooling to identify bottlenecks across layers for latency and throughput. Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. You might thrive in this role if you: Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. Have deep expertise with performance profiling, benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve o

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -83.9%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role The stakes of securing OpenAI increases as our internal coding and research becomes increasingly driven by autonomous AI agents. Compromising these agents could allow a cyber threat actor to compromise many other parts of the company. In this role, you would lead Preparedness work defending the security of our internal AI agents against insiders, Advanced Persistent Threats (APTs), or powerful AI agents. We’re looking for a strong hands-on technical executor with experience working directly with advanced cyber threat actors. In this role, you will: Develop and maintain threat models via which advanced attackers could compromise our coding assistants and automated security systems. Identify security investments that are especially critical to make in advance; for example, prioritizing by implementation lead-times, costs, and benefit. Partner with Security, Infrastructure, Research, Legal, and Preparedness to align on implementation plans and tradeoffs. Lead technical execution directly when needed, including prototyping controls, writing and reviewing software, and coordinating engineers across teams. Work with penetration testers to close gaps in defenses. You might thrive in this role if you: Are an exceptional hands-on technical executor. Have worked with advanced

AWSRestAIRust
🔔

Get new infrastructure team manager jobs in United States by email

Daily job updates · Unsubscribe anytime