Jobs in United States

Distributed Systems Engineer in San Francisco

130 active opportunities · Updated October 2026

Explore current distributed systems engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. About the Role The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity. In this role, you will: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output. Use our latest AI tools to re-think how we can be the most productive team in the industry. Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements. Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations. Guide and advise product engineering teams on best practices for ensuring observable, scalable systems. Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers. Have experi

PythonAWSKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in AWS-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including AWS-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re hiring Machine Learning Engineers to build and improve the AI systems that help strategic partners adapt OpenAI models to important use cases in cloud-native environments. This role spans post-training workflows, evaluation, data pipelines, model behavior, and API/infrastructure integration. You’ll work at the boundary between partner needs and core ML systems: helping teams understand what is and isn’t working, diagnosing issues in training and evaluation workflows, and turning those learnings into improvements to the underlying platform. You should enjoy working with external technical partners, extracting the real goal from messy requests, and pushing back or reframing when the requested experiment is not the highest-leverage path. You’ll collaborate closely with Research, Applied, Safety Systems, infrastructure teams, and external technical partners to solve ambiguous model-performance problems. When you succeed, strategic partners and internal teams will be able to improve model behavior with confidence, driving measurable product improvements while the systems behind that work become more reliable, scalable, and effective over time. In this role, you will Partner with strategic customers and in

PythonAWSKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

Join the engineering teams that bring OpenAI’s ideas safely to the world! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic-testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU, storage, GPU, and network lifecycle management to drive efficiency, accountability and dynamic optimization of our resources. Implement fault-tolerant and resilient design

AWSKubernetesRestMicroservices
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the role: We're seeking a Data Engineer to take the lead in building our data pipelines and core tables for OpenAI. These pipelines are crucial for powering analyses, safety systems that guide business decisions, product growth, and prevent bad actors. If you're passionate about working with data and are eager to create solutions with significant impact, we'd love to hear from you. This role also provides the opportunity to collaborate closely with the researchers behind ChatGPT and help them train new models to deliver to users. As we continue our rapid growth, we value data-driven insights, and your contributions will play a pivotal role in our trajectory. Join us in shaping the future of OpenAI! In this role, you will: Design, build and manage our data pipelines, ensuring all user event data is seamlessly integrated into our data warehouse. Develop canonical datasets to track key product metrics including user growth, engagement, and revenue. Work collaboratively with various teams, including, Infrastructure, Data Science, Product, Marketing, Finance, and Research to understand their data needs and provide solutions. Implement robust and fault-tolerant systems for data ingestion and processing. Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear. Ensure the security, integrity, and compliance of data according to industry and company standards. You might thrive in this role if you: Have 3+ years of experience as a data engineer and 8+ years of any software engineering experience(including data engineering). Proficiency in at least one programming language commonl

PythonJavaAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We are looking for an experienced Research Engineer to work on retrieval & search problems across our API and ChatGPT. As the AI landscape has evolved over the last few years, retrieval & search have emerged as key use cases for our models, and we are investing in ensuring that we can offer these search-based product experiences for our users. You will be at the center of our retrieval & search efforts as a company, and the progress you drive here will reach millions of end users. In this role, you will: Work on retrieval & search algorithms and methodologies in close collaboration with our research team, including problems in such domains as document search, enterprise search, knowledge retrieval, and web-scale search. Deploy these search methodologies into production in both the API and ChatGPT to be used by millions of end users. Explore novel research topics in retrieval & search that may inform our product strategy in the medium and long term. Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world You might thrive in this role if you: Have extensive prior experience building and maintaining production machine learning systems. Have prior experience working with vector databases, search indices, or other data stores for search and retrieval use cases Have prior experience building and iterating on internet-scale search systems Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done Have the ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or de

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. This role is based in San Francisco, CA. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl

AWSKubernetesRestAI
M
📍 San Francisco, United States· Full-time
✓ High-confidence listingCompany trend -55.6%
Quick readStrong listing-quality and freshness signals

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About Mixpanel Mixpanel turns data clarity into innovation. Trusted by more than 29,000 companies, including Workday, Pinterest, LG, and Rakuten Viber, Mixpanel’s AI-first digital analytics help teams accelerate adoption, improve retention, and ship with confidence. Powering this is an industry-leading platform that combines product and web analytics, session replay, experimentation, feature flags, and metric trees. Mixpanel delivers insights that customers trust. Visit mixpanel.com to learn more. About The Team Mixpanel Engineering is a small, fast-moving team focused on delivering real value to customers. We build powerful AI-powered product analytics while obsessing over clarity, simplicity, and delight. Engineers here own problems end to end. You can move across the stack to ship impact without being blocked by silos or heavy process. Product innovation drives our business, and product engineering teams own that responsibility. Our OLAP engine queries over 500 trillion events; a typical blob storage system we interact with processes 300 PiB/month at 1.2 Tbps sustained, and we run many of them across the world. The Data Runtime team owns the data execution layer that powers every Mixpanel product. We ensure that every customer query runs fast, cheap, and reliably, at any scale. This is an exciting time to join. Mixpanel's agentic and AI-first products are driving rapid growth in query volume, and Data Runtime is making the big bets that power it. We’re investing in elastic query compute and a distributed file cache that will let us scale query workloads dramatically without scaling cost with them. We

PythonSQLAWSAzure
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse, misalignment, and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. Within Safety Systems, the Model Policy team works to ensure that frontier models behave safely and reliably in real-world environments by designing policies that define safe model behavior. Some of our publications include: Safety at every step OpenAI GPT6 System Card OpenAI Model Spec About the Role We’re hiring a Model Policy Manager to shape model behavior for U.S. government use, with a focus on national security applications. You’ll define nuanced policies and translate them into training and evaluation criteria, helping models navigate high-stakes scenarios while preserving their usefulness and capabilities. In this role, you will: Develop model policies that guide safe and useful behavior. Build evaluations, identify policy gaps and model failures, and use findings to improve policies and training. Work with research, engineering, and domain experts to support safe, reliable deployment. You might thrive in this role if you: Bring relevant experience in AI safety, policy, or risk assessment. Have strong judgment and can turn complex safety questions into clear, practical policies. Have the technical fluency to work hands-on with model data and evaluations. Are motivated by OpenAI’s mission and the responsible use of

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse, misalignment, and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. Within Safety Systems, the Model Policy team works to ensure that frontier models behave safely and reliably in real-world environments by designing policies that define safe model behavior. Our relevant publications include: Safety at every step OpenAI GPT6 System Card OpenAI Model Spec GPT-Live ChatGPT Images 2.5 About the Role We are hiring a Model Policy Manager to focus on the safety of multimodal models. In this role, you will shape how OpenAI identifies, evaluates, and addresses risks in multimodal AI models - such as GPT-Live and ChatGPT Images - as well as multimodal capabilities in frontier AI models. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and maintain model policies for audio, image, video, and omni-modal behavior. Translate theories of harm and threat models into behavioral safety policies, evaluation criteria, grading guidance, and safeguards. Identify and analyze safety regressions and failure patterns to identify gaps in existing policies and inform policy iteration. Develop policy artifacts that support model training, evaluation, and deployment, including behavior i

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to train nuanced safety behaviors, how to make the model robust to bad actors, how to address privacy and security risks, and how to make the model trustworthy in safety-critical situations. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role We’re seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You’ll advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. In this role, you will: Research and implement methods for safety training, reinforcement learning, and adversarial robustness. Develop evaluations, identify model failure modes, and use findings to improve training. Work with research, engineering, security, and policy partners to support safe, reliable deployment. You might thrive in this role if you: Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness. Have a degree in computer science, machine learning, or a related field, and strong deep learning research or engineering skills. Have experience improving model safety for deployment and enjoy collaborative research. Are motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings. Security Requirements Active TS/SCI clearance or equivalent. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefi

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse, misalignment, and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. About the Role We are hiring a Product Manager to focus on risk related to multimodal models. In this role, you will drive initiatives which ensure that OpenAI’s audio, image, and video deployments are safe, impactful, and aligned with user needs and technical innovation. You will clarify strategic priorities, develop safety-focused product roadmaps, and collaborate closely with AI researchers, software engineers, policy experts, and cross-functional partners. This role suits a proactive, technically skilled product manager adept at adversarial thinking and excited to tackle challenging, ambiguous problems through structured analysis and collaborative decision-making. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Partner closely with AI research, engineering, data science, policy teams, and other stakeholders to embed safety throughout the development and deployment of multimodal AI models - such as GPT-Live and ChatGPT Images - as well as multimodal capabilities in frontier AI models. Develop comprehensive frameworks for understanding and mitigating deployment safety risks, drawing on data analysis, expert consultation, and adversarial assessments. Define strate

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Safety Systems manages the complete lifecycle of safety efforts for OpenAI’s frontier models, ensuring our models are deployed responsibly and have a positive impact on society. Our work spans diverse research and engineering initiatives—from system-level safeguards and model training to evaluation and red-teaming—all aimed at mitigating misuse and maintaining our high bar for safety. We lead OpenAI's commitment to developing and deploying safe Artificial General Intelligence (AGI), fostering a culture of trust, responsibility, and transparency. Our goal is to continuously learn from deployments, distribute AI’s benefits widely, and ensure that powerful tools remain aligned with human values and safety considerations. About the Role The Safety Measurement Product Manager owns OpenAI's approach to measuring harm and safeguard efficacy in production, including driving the strategy for our suite of safety measurement platforms and products used across the company. You will partner closely with our safety research and engineering teams to determine what we measure, where we measure it, and how we measure it, feeding those insights directly into critical leadership decisions and back into our safety work. You will also represent the company's topline safety metric as well as prioritize incoming requests from partner teams to expand our safety measurement platform to more use cases. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Partner closely with data science, research, engineering, policy teams, and other stakeholders to craft a vision for understanding safety outcomes and prevalence on our platforms. Define strategic priorities and product roadmaps focused on improving safety measurement approaches will scaling our measurement platform to more use cases, products, and cross-functional team needs. Establish repeatable processes to integrate cutting-edge AI safety research into OpenAI’s safety m

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Full-Stack Engineer to join our Data Acquisition team to build and optimize the interfaces and tools that power our data infrastructure. Responsibilities: Develop and maintain full-stack applications that support data acquisition, including internal tools and dashboards. Collaborate closely with cross-functional teams, including Data Processing, Architecture, and Scaling, to ensure seamless data ingestion and workflow management. Design and implement APIs to facilitate data interactions between internal services and external data sources. Enhance user experience by developing intuitive web-based interfaces for managing and monitoring data pipelines. Optimize backend services for performance, scalability, and security in a distributed computing environment. Work with legal and compliance teams to ensure our data acquisition processes adhere to privacy regulations and best practices. Deploy and maintain infrastructure using Kubernetes and Infrastructure-as-Code (IaC) methodologies. Analyze system performance, conduct experiments, and improve data workflows to maximize efficiency. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in full-stack development. Proficiency in frontend frameworks (React, Vue, or similar) and backend technologies such as Python, Node.js, or Go. Strong expertise in RESTful APIs, GraphQL, and database design (SQL and NoSQL). Experience building data-intensive applications that handle large-scale datasets. Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker). Prior experience with web crawling and large-scale data processing is a

PythonReactNode.jsVue
🔔

Get new distributed systems engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime