Jobs in United States

Back End Td Reliability Lab Manager in United States

490 active opportunities · Updated October 2026

Explore current back end td reliability lab manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team The Applied team safely brings OpenAI's technology to the world. We released ChatGPT; Plugins; DALL·E; and the APIs for GPT-5, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. Our customers build fast-growing businesses around our APIs, which power product features that were never before possible. ChatGPT is a prime example of what is currently possible. We simultaneously ensure that our powerful tools are used responsibly. Safe deployment is more important to us than unfettered growth. The Fraud Engineering team works within our Applied Engineering organization identifying and responding to fraudsters on our platform. We are looking for a software engineer with anti fraud & abuse experience to help architect and build our next-generation anti-fraud systems. About the role The Scaled Abuse team protects OpenAI’s products and customers by detecting, preventing, and responding to fraudulent and abusive behavior at scale. We build and operate the backend and data systems that power real-time detection, investigation workflows, and enforcement — balancing strong protections with a great user experience as the platform grows. Our work sits at the intersection of engineering and abuse expertise: we partner closely with Trust & Safety, Security, and Product to understand emerging attack patterns, translate messy signals into clear system behavior, and continuously harden our defenses. The problems are dynamic and ambiguous by default, so we value engineers who can quickly dive into an unfamiliar codebase, develop strong intuition about how it works end-to-end, and propose pragmatic improvements that make the entire stack more resilient. In this role, you will: Design and build systems for fraud detection and remediation while balancing fraud loss, cost of implementation, and customer experience Work closely with finance, security, product, research, and trust & safety ope

PythonAWSAzureKubernetes
O
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s API Platform organization builds the products and infrastructure that help first-party and third-party developers build with OpenAI models. We ship the API primitives, tools, SDKs, documentation, playgrounds, and platform experiences that make OpenAI’s capabilities reliable, understandable, and useful in production. The API Experience team is focused on the end-to-end developer experience for the OpenAI API. We own the surfaces developers touch every day: docs, SDKs, the Playground, examples, onboarding flows, and the systems that help developers go from first request to production deployment quickly and confidently. About the Role We’re looking for full stack and frontend engineers to help define and build the next generation of OpenAI’s developer experience. In this role, you’ll work across frontend product surfaces, backend systems, SDK and documentation pipelines, and API workflows that serve millions of developers and companies. You’ll partner closely with product, design, research, API engineering, and developer-facing teams to make complex AI capabilities simple to understand, easy to test, and safe to launch in real-world applications. This is a highly cross-functional role for someone who cares deeply about craft, developer empathy, reliability, and product velocity. In this role, you will: Build and scale developer-facing products including the OpenAI API Playground, documentation experiences, onboarding flows, examples, and API workflow tools. Own full stack projects end to end, from product definition and UX collaboration through backend implementation, launch, measurement, and iteration. Improve the systems that generate, maintain, and publish SDKs, API references, docs, guides, and developer examples. Partner with API, research, design, and infrastructure teams to bring new model capabilities and API primitives to developers in a clear, usable way. Use developer feedback, product analytics, and direct customer insight to identif

TypeScriptPythonReactAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

The ChatGPT Finances team builds experiences that help people connect their financial accounts, understand their financial picture, and ask useful questions about their finances through ChatGPT. Our work spans account connectivity, data ingestion, dashboards, personalized insights, and conversational experiences. We collaborate across product, design, research, infrastructure, security, and data integrations to make complex financial information understandable and actionable. This is an early and ambitious product area with a substantial roadmap. We are looking for engineers who want to shape both the first user experiences and the durable systems required to earn and keep users’ trust. About the role We’re looking for full-stack product engineers to build and scale ChatGPT Finances. You will own features across the stack—from polished frontend experiences to the APIs, services, and data models that power them. This role is well suited to engineers who combine strong product judgment with broad technical depth. You should care about how quickly users can understand their financial lives, how reliably data moves through the system, and how AI can answer financial questions in a grounded, transparent, and useful way. You will work closely with product, design, research, infrastructure, security, and data integration teams to take ideas from early prototypes to reliable production experiences. In this role, you will Own full-stack product features from user experience and frontend implementation through backend services, data models, deployment, and observability. Build polished, accessible, and performant interfaces for account connection, dashboards, insights, and conversational workflows. Design APIs and backend systems that safely ingest, normalize, and serve financial data. Build resilient integrations that handle synchronization, data freshness, partial failures, permissions, and user consent. Bring new AI capabilities into production while prioritizing grounding

TypeScriptPythonReactNode.js
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects, and spend less time on repetitive tasks. AI has already transformed how code is written, but software engineering extends far beyond coding. Our mission is to apply AI across the entire software development lifecycle (SDLC) — from design and implementation to code review, testing, debugging, issue remediation, maintenance, documentation, and user support. The team is also responsible for developer-facing Codex experiences including the Codex IDE Extension and the terminal interface, which are used daily by developers ranging from individual open-source contributors to some of the world’s largest engineering organizations. The team also works closely with the open-source software community, building tools that help maintainers and contributors manage increasingly complex projects. We believe AI can make open-source development more sustainable by reducing the operational burden of reviewing contributions, triaging issues, maintaining quality, and supporting growing communities. By building the future of software development, we're helping advance OpenAI's mission of ensuring that the benefits of AI reach people around the world. About the Role We’re hiring a Full Stack Software Engineer to help invent the next generation of AI-powered software development workflows. “Full stack” in this role means much more than traditional frontend and backend development. You'll own complete product experiences, spanning user interfaces, workflow orchestration, agent and prompt design, backend systems, and cloud infrastructure. This is a highly product-oriented role. You'll work directly on the workflows developers use every day, identifying bottlenecks and rethinking how software gets built in a world where AI agents are active participants in the development process. The features you ship will inf

TypeScriptAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Cybersecurity Products team builds products at the frontier of AI and cybersecurity. Our work includes Codex Security and related cyber products that turn advances in model capability into dependable tools for defenders. We help teams find, validate, and remediate vulnerabilities, continuously improve the security of software, and test AI-powered applications before they reach production. About the Role As a Full Stack Software Engineer, you will build the product experiences and systems that make AI-powered security useful in real engineering environments. You will work across web surfaces, APIs, orchestration, data models, and integrations to help security and engineering teams move from a codebase or application to evidence-backed findings, prioritized remediation, and revalidation. You will collaborate closely with product engineers, security researchers, and customer-facing teams. The work spans fast-moving product development and hard systems problems: long-running workflows, large repositories, sensitive data, reliability, observability, and a high bar for earning user trust. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build end-to-end workflows for vulnerability discovery, security scanning, red teaming, findings review, remediation, and reruns. Design and operate backend services for long-running security work, including APIs, asynchronous orchestration, durable state, and integrations with developer workflows. Make complex security results actionable through clear product surfaces, strong evidence, thoughtful prioritization, and reliable reporting. Partner with security researchers, product teams, and users to evaluate quality, reduce noise, improve coverage, and ship safely. You might thrive in this role if you: Have experience shipping production full-stack products across modern web frontends and backend s

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The ChatGPT organization at OpenAI supports our mission by bringing advanced AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent breakthroughs in multimodal AI have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering. We're building the systems and experiences that turn these research advances into products used daily by creators, professionals, businesses, and consumers around the world. Our team sits at the intersection of research, product, design, and infrastructure. We work closely with model researchers, mobile engineers, frontend engineers, and platform teams to build intuitive experiences and scalable systems that power image generation at global scale. Whether users are creating marketing assets, visualizing ideas, editing photos, designing products, or simply exploring their creativity, our goal is to make visual creation feel as natural as having a conversation. About the Role We are looking for an experienced Full Stack Engineer to join the Image Generation team and help shape the future of AI-powered visual creation. In this role, you'll own features end-to-end across both frontend and backend systems, building the experiences that enable users to generate, edit, organize, and interact with images inside ChatGPT. You'll work across the entire stack—from highly interactive user interfaces and real-time workflows to backend services, APIs, orchestration systems, and data infrastructure. This role is ideal for engineers who enjoy moving fluidly between product development and systems engineering, collaborating closely with design, product, and research teams to rapidly bring new AI capabilities to users. You'll help define entirely new interaction paradigms as multimodal AI continues to evolve. In

TypeScriptReactAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Applied Foundations team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Applied Foundations team is at the front lines of defending against financial abuse, scaled attacks, and other forms of misuse that could undermine the user experience or harm our operational stability. Integrity Foundations provides the core building blocks and infrastructure for this work. About the Role At OpenAI, our mission is to advance AI in a way that is safe, reliable, and aligned with broad societal values. The applied foundations role is crucial for maintaining the trustworthiness of our platforms. You will be pivotal in developing robust defenses against a spectrum of adversarial behaviors that threaten our ecosystem. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. In this role, you will: Develop and enhance systems to detect and prevent various forms of abuse including financial fraud, botting, and scripting. Collaborate with cross-functional teams to design solutions that protect against and mitigate adversarial attacks without compromising user experience. Assist with response to active incidents on the platform and build new tooling and infrastructure that address the fundamental problems. You might thrive in this role if you: Have at least 3 years of professional software engineering experience. Have experience setting up and maintaining production backend services and data pipelines. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Are self-directed

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i

PythonReactNode.jsAngular
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Future of Computing Research team is an Applied Research team within the Consumer Devices group focused on developing new methods and models as we advance forward in our mission of building AGI that benefits all of humanity. As a Software Engineer on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. About the Role We are looking for a Software Engineer to join our team to build tools and services that enable AI research, evaluation, and data generation workflows. The best work in this role will start with an ambiguous design question and turn it into working research systems. You will work closely with researchers, designers, and engineers to build the evaluation systems, synthetic data generation pipelines, review tools, and supporting platform services. The goal is to make these workflows easier to create, run, and trust without requiring bespoke engineering support for each new design concept. You will help ensure that research artifacts have a clear lifecycle, runs are reproducible and observable, and results provide useful evidence for product and model-training decisions while the underlying systems remain reliable and reusable. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Build web applications, APIs, data models, and backend services for AI research workflows. Build tools to author and manage evaluation tasks, rubrics, graders, suites, and rollout configurations, including workflows for publishing, versioning, auditing, and sharing research artifacts. Automate evaluation runs and generate useful reports for design, research, and engineering teams. Support synthetic data generation workflows for multimodal and conversational research, including tools that comb

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The ChatGPT team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn from deployment and broadly distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. We aim to make our innovative tools globally accessible, transcending geographic, economic, or platform barriers. Our commitment is to facilitate the use of AI to enhance lives, fostered by rigorous insights into how people use our products. About the Role We are looking for an experienced fullstack engineer to join our new ChatGPT Growth team to spearhead high-impact projects that amplify the user base of ChatGPT and Plus Subscribers. Your role will include projects such as optimizing account access, notifications, SEO, fostering value discovery, and virality. As we are in the nascent stages of growth at OpenAI, we will rely on you to discover pivotal areas where strategic bets or incremental efforts can catalyze significant impact. We value engineers who are impact-driven, autonomous, adept at discerning crucial insights from experimental results, and have a strong intuition for how to remove barriers to unlocking the magic of ChatGPT. In this role, you will: Drive long-term growth of ChatGPT through a combination of data analysis, product ideation, and experimentation to optimize product experiences. Plan and deploy backend APIs necessary to power these product experiences. Execute on projects by working closely with research, product, design, data science and other members of product teams to land impact on product goals. Create a diverse and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group-think. You might thrive in this role if you: Shipped features on web that optimize the user funnel, such as landing pages, product pages, purchase flows, search flows, etc. Are highly analytical and have experience designing and implementing A/B te

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to, or even ask to take on future tasks proactively. Our team is a fast-moving group within OpenAI, bringing together research, engineering, design, and product. We iteratively build the Codex agent harness and product to get the most out of the model, and we iteratively train the model to be great at complex software engineering tasks. The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. We operate across research, engineering, product, and infrastructure; owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. Codex Enterprise builds the ecosystem, governance, and enterprise capabilities that help Codex spread across developers, teams, and organizations worldwide. The Enterprise Controls team owns the systems that allow companies to safely deploy Codex across their organization while protecting their most sensitive code, data, and internal knowledge. About the Role As Codex adoption grows inside large organizations, customers are increasingly trusting Codex with their most valuable assets: proprietary codebases, internal documentation, customer data, and sensitive workflows. This role will help build the enterprise control plane that makes Codex secure, governable, and trustworthy at scale. You will design and operate backend systems that give enterprise administrators visibility and control over how Codex is used across their organization. You will work across identity, access, encryption, policy enforcement, auditability, and admin controls. This may include systems that let customers manage encryption keys, control which Codex capabilities are enabled, enforce organizational policies, and understand how data flows through Codex. This role owns systems end-to-end: from architecture and

PythonJavaAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. About the Role We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. Design and iterate on user-facing tools and backend services to support high-quality data workflows Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal r

AWSRestAIRust
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team Online Data builds and operates Habitat, the single product surface of Online Data and the system of record for OpenAI’s online user data. As OpenAI’s scale and product requirements evolve, Habitat is becoming a full-stack, one-size-fits-most database platform with end-to-end ownership of: Provisioning and developer experience APIs and guardrails Scaling, performance, and reliability Data movement, caching, routing, and placement Privacy enforcement and access control Change Data Capture (CDC) as a first-class primitive The foundation for future storage backends You’ll work on the core online database platform behind OpenAI’s products, building and operating Habitat services that handle high-QPS, latency-sensitive workloads across regions. You’ll partner closely with internal platform and product teams to ship safe, reliable systems, then push them to be faster and more cost-efficient through better caching, routing, observability, and operational tooling. This is a critical role for engineers who like owning hard distributed-systems problems end to end and sweating the details from p99 latency to production operations at massive scale. In this role, you will Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement Own a major surface area end to end, from product and API design to operational excellence Improve latency, correctness, and cost efficiency for real production workloads at massive scale Build strong instrumentation, debugging workflows, and developer-first tooling Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability You might thrive in this role if you have A strong track record building and operating high-scale backend or data-intensive distributed systems in production Excellent systems judgment and the a

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior Android engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable Android foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Have a proven track record of building high-quality Android applications in production. Are fluent in Kotlin (and/or Java) and familiar with Android development tools and architecture components. Prioritize performance, security, and user experience in mobile development. Enjoy working cross-functionally to bring ambitious product ideas to life. Care deeply about performance, security, and user experience. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We

JavaAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the core infrastructure behind OpenAI’s monetization and ads systems. In this foundational role, you’ll architect and implement distributed systems that power OpenAI’s monetization stack—focusing on reliability, performance, privacy, and large-scale operation. You’ll work across backend, systems, and platform layers to define and implement 0→1 infrastructure, partnering closely with Product, Design, and Research to shape the future of monetized AI experiences. Your work will enable both internal and external teams to build on safe, scalable, and robust monetization primitives. This role is exclusively based across our San Francisco & Seattles sites. We offer relocation assistance to new employees. In this role, you will: Design and build the foundational backend and infrastructure powering OpenAI’s monetization and ads systems Architect large-scale distributed systems that

AWSRestAIGo
🔔

Get new back end td reliability lab manager jobs in United States by email

Daily job updates · Unsubscribe anytime