About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
Jobs in United States
Aws in San Francisco
866 active opportunities · Updated October 2026
Showing
15 jobs
Explore current aws jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Safety Systems team is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society, and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Safety Oversight Research team aims to fundamentally advance our capabilities to maintain oversight over frontier AI models, and leverage these advances to ensure OpenAI’s deployed models are safe and beneficial. This requires a breadth of new ML research in the areas of human-AI collaboration, reasoning, robustness, and scalable oversight to keep pace with model capabilities. We invest heavily in developing novel model and system-level methods of identifying and mitigating AI misuse and misalignment. Our goal is to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role OpenAI is seeking a senior researcher with a passion for AI safety and experience in safety research. Your role will set directions for research to maintain effective oversight of safe AGI and work on research projects to identify and mitigate misuse and misalignment in our AI systems. You will play a critical role in defining how a safe AI system should look in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI. In this role, you will: Develop and refine AI monitor models to detect and mitigate known and emerging patterns of misuse and misalignment. Set research directions and strategies to make our AI systems safer, more aligned, and more robust. Evaluate and design effective red-teaming pipelines to examine the end-to-end robustness of our safety systems, and identify areas for future improvement. Conduct research to improve models’ ability to reason about questions of human values, and apply these improved models to practical safety challen
About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Strategic Sourcing & Procurement function plays a critical role in enabling OpenAI to deliver impact across research, product development, technology infrastructure, and services by helping the company scale responsibly, securely, and with strong commercial discipline. Our work sits at the intersection of innovation and execution. We partner closely with teams across OpenAI to translate rapidly evolving business needs into scalable, compliant, and economically sound external partnerships. As OpenAI continues to grow at pace, services sourcing is becoming increasingly strategic across the company. Every business unit relies on external service providers in different ways — to extend capacity, access specialized expertise, support operations, and accelerate execution. Done well, Procurement becomes a source of trust and momentum, helping OpenAI move faster with the right partners, stronger commercial outcomes, and the right level of protection. About the Role We are seeking an experienced Strategic Sourcing (GTM) Leader to lead strategic sourcing and commercial enablement for OpenAI’s Go-to-Market organization across B2B and B2C channels. You will manage substantial and rapidly growing spend while shaping sourcing strategies and scalable commercial pathways across Media, Creative, Production, Influencer, Agency, Sponsorships, Analytics, Communications, and Event suppliers in support of high-impact global initiatives. You’ll help evolve our GTM procurement function from reactive deal support into a speed-enabling, scalable commercial engine that delivers cost efficiency, launch readiness, and strong governance in a fast-moving environment. In this role, you will: Develop and execute sourcing strategies across GTM, Brand, Global Affairs, Events, Growth, and Partnership activities—spanning both B2B and B2C channels—that align with our mission and b
AI Systems Engineer - Codex Core Agents About The Team The Codex Core Agents team builds the agent harness that turns model capability into real-world action. We own the systems around the model: prompting and interpreting model outputs, executing actions safely in real environments, and feeding production experience back into better models and better agent behavior. This team sits close to research and works across the stack: harness, model interaction, inference, sandboxed execution, orchestration, evals, production reliability, and the performance envelope around tokens, latency, cost, capacity, and quality. The harness is open source and increasingly part of how models are trained and evaluated, making this one of the highest-leverage layers in Codex. About The Role We’re looking for engineers to build the AI systems that make Codex agents dependable in production. The ideal candidate is an agent-systems builder: hands-on across low-level systems and ML workflows, able to debug Codex behavior end to end across the harness, model behavior, inference/runtime stack, GPU fleet, and product surface. You’ll work with research, infrastructure, and product to design agent harness capabilities, run experiments and ablations across the model + system prompt + harness stack, build frameworks for assessing production agent performance, and turn messy failures into durable improvements. What You’ll Do Design and build the core agent harness and execution loop that lets Codex agents interpret model outputs, use tools, execute code, and complete long-horizon tasks safely. Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents operating in real development environments. Develop evaluation, experimentation, and debugging systems that distinguish harness issues, model behavior, inference/runtime issues, and product failures. Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to
About the Team The Youth Well-Being product team is part of the Integrity pillar at OpenAI, responsible for ensuring that our state-of-the-art AI technologies are deployed in safe, age-appropriate, and beneficial ways—especially for youth and families. We work across OpenAI’s entire product surface area, from ChatGPT to future-facing tools, to architect safety and trust into the foundation of our systems. Our mission: empower families while meeting the highest standards of regulatory compliance and ethical responsibility. This work is core to OpenAI’s mission to ensure AGI benefits all of humanity. Safety, especially for the most vulnerable users, is more important to us than unfettered growth. About the Role We’re looking for a Senior Software Engineer to help architect and build the foundational systems that power family- and youth-facing experiences at scale. You’ll help define how teens and guardians engage with OpenAI products—ensuring those experiences are safe, compliant, and empowering. You’ll operate across a broad technical surface: building identity primitives, age assurance pipelines, and guardian tools that support both proactive and reactive interventions. You’ll collaborate closely with a cross-functional team of engineers, data scientists, designers, user researchers, and policy experts. This is a high-impact, 0→1 opportunity to set the standard for how families interact with generative AI. In this role, you will: Architect and implement teen and guardian experiences across OpenAI products, including ChatGPT. Build global age assurance systems that are privacy-preserving and tailored to regional compliance needs. Design and evolve our identity infrastructure to support scalable, secure, and resilient user journeys at consumer internet scale. Help define safety and well-being metrics, and continuously improve user trust through technical interventions. You might thrive in this role if you: Have built products or infrastructure for users under 18, or h
About the Team The Interpretability team studies internal representations of deep learning models. We are interested in using representations to understand model behavior, and in engineering models to have more understandable representations. We are particularly interested in applying our understanding to ensure the safety of powerful AI systems. Our working style is collaborative and curiosity-driven. About the Role OpenAI is seeking a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, in close collaboration with a highly motivated team. You will play a critical role in helping OpenAI ensure future models remain safe even as they grow in capability. This will make a significant impact on our goal of building and deploying safe AGI. In this role, you will: Develop and publish research on techniques for understanding representations of deep networks. Engineer infrastructure for studying model internals at scale. Collaborate across teams to work on projects that OpenAI is uniquely suited to pursue. Guide research directions toward demonstrable usefulness and/or long-term scalability. You might thrive in this role if you: Are excited about OpenAI’s mission of ensuring AGI benefits all of humanity, and are aligned with OpenAI’s charter . Show enthusiasm for long-term AI safety, and have thought deeply about technical paths to safe AGI. Bring experience in the field of AI safety, mechanistic interpretability, or spiritually related disciplines. Hold a Ph.D. or have research experience in computer science, machine learning, or a related field. Thrive in environments involving large-scale AI systems, and are excited to make use of OpenAI’s unique resources in this area. Possess 2+ years of research engineering experience and proficiency in Python or similar languages. Are deeply curious. About OpenA
About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We’re looking for dedicated, experienced, and deeply curious individuals to help solve some of the most complex challenges faced by our customers while building the future of post-AGI support alongside us. In this role, you’ll work directly with customers through support tickets and live calls, troubleshooting high-impact issues and resolving novel, often ambiguous technical problems in one of the fastest-moving environments in technology. As AI adoption rapidly accelerates, the work you do will directly support mission-critical systems being built on OpenAI’s platform, serving as a critical line of defense for customers operating at massive scale. Beyond resolving technical issues, you’ll help define what world-class support looks like in an AGI-driven future. You’ll partner closely with Engineering, Product, and Operations to improve systems, reduce bugs, and elevate the customer experience, while leveraging automation, agents, and our own AI technology to transform how support operates at scale. You’ll be responsible for: Working directly with customers to troubleshoot and resolve their most complex technical issues, including API failures, integration challenges, authentication errors, and production incidents. Providing end-to-end ownership through debugging logs, analyzing system behavior, reproducing issues, and
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b
About the Team The Core Services team is responsible for building and managing foundational services. It acts as the bridge between core infrastructure (e.g. compute, storage, networking) and product engineering teams, and enables product teams to move fast, build reliably, and scale efficiently. About the Role As a software engineer in the core services team, you will design and operate critical backend platforms such as caching systems, workflow orchestration, metadata stores, and file services. You’ll focus on building highly reliable, scalable, and performant systems that serve as the backbone of our products. We’re looking for people who are passionate about building infrastructure that empowers product teams, love working on distributed systems challenges, and enjoy creating well-designed APIs and abstractions that accelerate development. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain shared infrastructure services such as caching layers, workflow orchestration (Temporal), metadata stores, and file storage services. Collaborate with product teams to provide scalable, reliable primitives that abstract the complexities of distributed systems. Improve performance, resilience, and scalability of core services that power customer-facing applications. You might thrive in this role if you: Have experience with distributed systems, caching infrastructure (e.g., Redis, Memcached), metadata storage (e.g., FoundationDB), or workflow orchestration (e.g., Temporal, Cadence). Have experience running containerized services in cloud environments and integrating them into automated build/test/release (CI/CD) workflows. Understand trade-offs in consistency models, replication strategies, and performance optimization in multi-region systems. Excel at communication and collaboration with cross-functional teams, and are obsesse
About the Team The Fleet team builds core components to enable productive research from small to state of the art scale across OpenAI, with the goal of accelerating progress towards AGI. We frequently collaborate with other teams to speed up the development of new state-of-the-art capabilities. About the Role As we scale up with more researchers and engineers joining OpenAI, we seek a pragmatic and passionate engineer with a strong focus on the development experience for both engineers and scientists. In this role, you will be responsible for building and maintaining systems that allow our research + engineering organization to iteratively develop, test, and deploy new features reliably, with high velocity, and with a frictionless and fast development cycle. You will help oversee and drive to the vision of how we should build, test and deploy software. You will drive the design of our continuous integration pipelines, testing infrastructure, training and support around our build system. Our current environment relies heavily on Python, Rust, and C++, which you will take ownership of and strive to transform into a state of the art development experience for research. Ultimately, your role will be to provide the necessary tools and metrics to support our fast-paced culture and ensure a stable, scalable platform for growth, while also fostering a seamless and low friction experience for OpenAI’s research. This role is based in San Francisco, CA. For a San Francisco role, we use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have supported large monorepo development and deployment before Are a proficient Python programmer working in large monorepos Are proficient with Docker and Kubernetes Experienced in CI/CD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boun
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil
About the Team We’re hiring software engineers to make the Workload team more productive. The Workload team maintains the core components of OpenAI’s training and inference frameworks and helps execute frontier experiments. About the Role We’re looking for someone who cares about the developer experience of working in and around OpenAI’s core training and inference frameworks. In this role you will: Be responsible for optimizing the development workflows of the engineers around you Work within various Workload teams to address their specific needs, but collaborate with the centralized teams that own various aspects of development experience Optimize iteration speed, both broadly, and in particular by optimizing specific teams’ CI Improve reliability, for instance, by driving testing strategy for particular components Work through the long tail of things that it takes to build libraries and systems that will delight researchers You might thrive in this role if: You are motivated by helping people. You believe a thing that separates great teams from good teams are the players willing to do whatever work it takes, without ego. You believe in the power of developer experience. Something magical happens when people can quickly and confidently iterate on a simple codebase, but this magic is fragile and must be fought for. When you see someone trip over something, no matter how small, your first instinct is asking yourself what it would take for that to not happen again. Your second instinct is clicking merge on the PR you’ve already written to make it so. You are pragmatic. You have the ability to see the world through a perfectionist’s eyes, but are not yourself a perfectionist. You know which problems to pick and when to switch to making progress on a different problem. You like going end-to-end on things. You love co-design — that feeling when you were only able to find the right solution because you both deeply understand the users that interact with a system and the
About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building a model powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for a Backend Software Engineer to help architect and scale the infrastructure that powers our knowledge systems. This is a deeply technical and highly cross-functional role where you’ll build robust systems and backend services that serve as the foundation for how knowledge is created, accessed, and applied across OpenAI. In this role, you will: Design, build, and maintain backend services and APIs to support intelligent automation and knowledge systems Integrate and structure data across internal platforms, transforming it into formats optimized for use by downstream systems and AI workflows. Collaborate closely with product, research, and engineering teams to integrate OpenAI mode
About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr
About the Team: OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. Our Stargate program develops and deploys massive, state-of-the-art data center campuses in partnership with industry leaders today—and through future OpenAI infrastructure projects tomorrow. We design for scale, speed, and reliability, and we need experienced technicians who can translate network blueprints into physical reality. About the Role: We are seeking a Senior Data Center Networking Technician who thrives in fast-moving build environments and is eager to roll up their sleeves during active datacenter deployments. Your first assignment will focus on the physical bring-up of network infrastructure at a large partner-operated campus, collaborating with partner teams and their delivery vendors to achieve agreed performance and reliability targets. As that campus reaches steady state, you will transition to lead network deployment for future OpenAI data center projects, defining standards and guiding implementation across multiple locations. Candidates must be able to sit onsite in Abilene, Texas 5 days per week Key Responsibilities Serve as OpenAI’s technical lead technician during the current campus build, partnering with internal engineers and external contractors on design reviews, installation plans, and acceptance criteria. Spend significant time on the data-center floor performing inspections, assisting with cable routing/termination when needed, conducting fiber testing (OTDR, power levels, continuity), and resolving installation challenges in real time. Troubleshoot and optimize cabling routes, patching, and equipment turn-up to ensure clean, reliable handoff to network operations. Contribute to design discussions and peer reviews for structured cabling and physical network layouts, providing practical field feedback to engineering teams. Develop repeatable engineering standards, as-built do
Other cities to consider
More places hiring for this role
Get new aws jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime