Jobs in United States

Reliability Engineer in San Francisco

227 active opportunities · Updated October 2026

Explore current reliability engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

P
📍 San Francisco, CA, United States· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Sr. Manager to lead our Capacity Engineering team. The team ensures that Pinterest’s cloud infrastructure has the capacity it needs while operating reliably, efficiently and with clear financial accountability. You’ll lead the full portfolio across forecasting and supply, capacity-management systems, compute and GPU efficiency, infrastructure data and governance and capacity operations. What you’ll do: Lead the Capacity Engineering team and establish its 12–18 month functional and technical strategy, roadmap and success measures tied to Infrastructure and company goals. Develop CPU and GPU forecasts and supply plans that account for workload demand, delivery constraints, cost and reliability requirements. Guide the design and delivery of capacity requests, reservations, entitlements, allocation policy and infra

KubernetesAIFinance
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Consumer Monetization builds the experiences and systems that power how customers purchase and pay for OpenAI products. Our scope spans purchasing flows, payments, subscriptions, and billing, along with the shared capabilities that support new products, offers, and distribution channels. We own both customer-facing experiences and the underlying platforms that power them. We partner closely with Product, Design, Growth, Data Science, and engineering teams across OpenAI to make purchasing effective and reliable, and new offerings easier to launch and monetize. About the Role We’re looking for experienced Staff+ engineers to evolve the payments, billing, and subscription capabilities that support OpenAI’s growing product portfolio. You’ll tackle problems where correctness, reliability, and flexibility are essential: supporting new billing requirements, managing the billing lifecycle, synchronizing state across internal systems and external providers, and enabling new products and commercial models. Your work may span several areas based on your expertise and team priorities: Billing and monetization capabilities: Extend billing capabilities and improve integrations and state consistency across systems to support new products and business models. Subscriptions: Orchestrate purchases, renewals, plan changes, cancellations, and recovery, ensuring customers are charged correctly and receive the right benefits. Payments: Expand payment capabilities through processor integrations, routing, and broader payment-method coverage. Risk and integrity: Partner with risk and integrity teams to integrate controls into purchasing flows, reducing abuse while protecting legitimate customer experiences. You’ll help set technical direction while remaining hands-on in implementation and delivery. This is an opportunity to solve complex engineering problems at scale, connect architecture decisions to customer and business outcomes, and help other engineers take on broader ow

Artificial IntelligenceAIFinance
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Plugin Ecosystem team builds the platform and product experiences that let people extend ChatGPT and Codex. We work on plugins, skills, connectors, interactive apps, and open standards like the Model Context Protocol (MCP). We make plugins easy to discover, install, and use, ensure they’re invoked at the right time, and help people find new ways to get value from them. We want anyone to be able to turn a useful workflow into a plugin, share it, and have other people use it. A plugin can package instructions and skills with connections to the tools and data it needs. Our work spans creation and publishing, reliable execution across our products, clear permissions and approvals, and the controls admins need to bring plugins to their organizations. We work closely with research to improve plugin quality as models evolve. About the Role We’re looking for product-minded engineers to build the systems behind plugins and improve how models use them. Depending on your focus, you may scale generalist infrastructure and identity-related integrations across products, or improve plugin quality at the intersection of backend engineering and applied AI or work on the product experience itself to drive plugin usage. You’ll work across teams and own problems from diagnosis and design through implementation and release. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and ship APIs, SDKs, and services that developers use to extend ChatGPT and Codex. Build intuitive experiences that help users discover, install, and use plugins to get more done. Make plugins easier to create, test, publish, update, and share. Improve when and how models use plugins, from choosing the right plugin to completing a task. Work with Research to diagnose failures and measure improvements as models evolve. Improve plugin reliability and interaction quality acros

Artificial IntelligenceAI
N
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86%
Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion’s Data Foundations team builds and operates the batch and streaming infrastructure behind our product features, analytics, search, and AI experiences. We’re looking for a hands-on technical leader to shape the next generation of this platform as Notion serves larger customers, expands globally, and supports more data-intensive products. You’ll identify the highest-leverage problems, set direction, build and develop a high-performing team, and lead multi-quarter initiatives across our data lake, streaming, distributed-compute, governance, and reliability systems. You’ll stay close to critical technical decisions while creating clear ownership, growing engineers and technical leaders, and helping the team execute as one—partnering closely with Data Engineering, Data Product, Search, AI, Infrastructure, and Security. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You

P
📍 San Francisco, CA, United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%

From $1.5M/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Job Title: Software Engineer II, Data Analytics and Engineering Intro: We’re looking for a Software Engineer II, Data Analytics and Engineering to improve the quality, reliability and velocity of data science and product development at Pinterest. You’ll build scalable data foundations, analytics tooling and analysis pipelines that enable trusted, self-service access to datasets, insights and metric investigations across cross-functional teams. What you’ll do: Develop and document practical instrumentation and experimentation standards, then partner with product engineering teams to apply them to priority product development work. Build and improve scalable analysis pipelines and tooling that produce reliable insights at scale and strengthen understanding of key data structures and metrics. Create tools and processes that enable Data Scientists a

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. Within Premium Support, Dedicated Support Engineers combine deep technical troubleshooting with an enduring understanding of our most strategic customers’ architectures, critical workloads, and business priorities. Through proactive reliability work, ownership during incidents, and AI-powered support capabilities, we help customers operate successfully as their use of OpenAI grows. About the Role We’re looking for a senior leader to build and scale our Dedicated Support Engineering function globally. You will define its strategy, build the team, and establish how we deliver technically rigorous, proactive support for customers running some of the most complex and consequential workloads on OpenAI. This role combines organizational leadership, technical judgment, and executive customer engagement. You will establish a model in which DSEs develop deep customer context, independently advance difficult investigations, anticipate operational risks, and drive issues through resolution. You will also turn what the team learns into improvements that benefit customers across OpenAI. You should bring experience building technical organizations that maintain long-term accountability for enterprise customers. Leadership in Technical Account Management, enterprise Support Engineering, or a comparable technical customer function is particularly rel

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a product manufacturing & quality engineer, who will be responsible for driving technical initiatives related to the manufacturing, quality and reliability of our AI supercomputer hardware systems to ensure product success from concept to launch and through mass production. You’ll have the opportunity to coordinate with functional SMEs and work with a wide range of stakeholders, from design engineering and operations teams, TPMs, external industry vendors and partners to ensure that all products are developed and delivered on time and to the highest quality standards. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Own the integrated manufacturing and quality readiness for a product across L6, L10, and L11, with clear gates, milestones, deliverables, owners, and closure criteria. Lead readiness of process flows, tooling, fixtures, assembly operations, test interfaces, and production controls. Review and contribute to work instructions. Translate product requirements into qualification plans, process controls, test requirements and acceptance criteria with design engineering and Area SMEs Coordinate and drive execution of product and process qualification, reliability testing, and validation with the relevant SMEs. Maintain traceable evidence that assigned products and processes meet agreed performance, reliability,

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Applied organization brings OpenAI’s most advanced technology to the world through products like ChatGPT and the APIs that power a growing ecosystem of developer and enterprise applications. Data Engineering builds and operates the trustworthy, secure, and reliable data systems that power decisions across OpenAI. About the Role We’re looking for a Data Engineering Manager to lead the Growth & Revenue data engineering team. This leader will own the data strategy and execution for the data subject areas spanning growth accounting across all product surfaces, product partnerships, checkout, billing, payments, revenue, and monetization, helping OpenAI understand how people adopt, engage with, and pay for our products. You will partner closely with several Data Science, Business, and Engineering partners to connect product behavior to trustworthy subscriber, payment, and revenue measurement. In this role, you will: Build, manage, and grow a high-performing, inclusive team across the Growth & Revenue data subject areas. Define the data strategy for all the data subject areas you own. Deliver durable, well-modeled data products that connect product behavior, subscription state, checkout events, payment outcomes, and revenue. Establish trusted metric definitions and data quality standards so product, growth, finance, and executive leaders can make fast, consistent decisions. Partner with Data Science and Product teams to support experimentation, causal measurement, funnel analysis, and scalable self-serve analytics. Partner with Finance and Financial Engineering to ensure analytical revenue views reconcile to financial truth and production billing systems. Raise operational excellence for critical pipelines, including reliability, observability, privacy, governance, and incident response. Set a clear roadmap, make principled tradeoffs, and communicate progress and risk across technical and business stakeholders. You might thrive in this role if yo

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The ChatGPT organization at OpenAI supports our mission by building products that bring cutting-edge AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research and product, partnering closely with researchers, designers, product managers, and platform engineers to bring state-of-the-art image generation capabilities to life across ChatGPT and our mobile applications. Millions of users rely on these experiences every day to create, communicate, learn, and build. About the Role We are seeking an experienced Android Software Engineer to build and improve image generation experiences within the ChatGPT Android app. You will help define how users create, edit, and interact with visual content powered by the latest multimodal AI models. This is an opportunity to work on a highly visible product area, translating cutting-edge AI capabilities into intuitive, performant, and delightful mobile experiences used by millions around the world. ChatGPT's Android app already enables users to generate and transform images directly from their devices, and we're just getting started. In this role, you will: Build and ship new Android features that power image generation and image editing experiences. Create intuitive user experiences that make advanced AI capabilities feel seamless and accessible. Collaborate closely with Product, Design, Research, and Engineering teams to bring new multimodal capabilities to production. Drive improvements in app performance, reliability, architecture, testing, and developer tooling. Optimize media-heavy workfl

AWSRestAIKotlin
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI develops models that can reason through complex problems and hardware designed for the demands of advanced AI. AI for Chips connects these efforts: applying increasingly capable AI systems to the work of semiconductor engineering. Our goal is to help engineers develop better chips and shorten design cycles. This work brings research, model training, and hardware expertise together to build tools that engineers can use on real designs, with correctness and measurable performance at the center. About the Role We’re hiring a Software Engineer to build the research infrastructure and tooling that help OpenAI models design silicon. You’ll turn chip-design workflows into reliable environments for reinforcement learning and evaluation, and make it easier for researchers to run experiments and iterate on new ideas. You’ll move between software engineering, tool integration, and open research problems. We value strong coding fundamentals, clear technical judgment, and independent execution. Prior chip-design experience is helpful, but you can learn the domain alongside the team’s hardware specialists. In this role, you will: Build and maintain infrastructure for reinforcement learning environments, evaluations, and long-running experiments. Integrate electronic design automation (EDA) tools into workflows for RTL generation, verification, and physical design optimization. Improve experiment reliability, reproducibility, observability, and performance; debug failures across tools, services, and infrastructure. Develop tooling and model harnesses that let researchers test ideas quickly and measure correctness and power, performance, and area (PPA). Collaborate with researchers and engineers to turn successful experiments into reusable systems and training workflows. Own ambiguous projects end to end, communicate progress, and use results to guide the next iteration. You might thrive in this role if you: Have strong software engineering fundamentals, with

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Safety Systems org is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Safety Engineering team builds the platforms and tools that make OpenAI’s models safe to use in the real world. We partner closely with researchers, product teams, and policy to turn safety ideas into reliable, scalable systems: measuring risk, enforcing safeguards, and continuously improving how models behave in production. Our work sits at the intersection of product engineering, data, and AI, and directly shapes how millions of people experience OpenAI’s technology. About the Role We’re looking for a self-starter engineer who loves building products in an iterative, fast-moving environment—especially internal tools that unlock real-world impact. In this role, you’ll build full-stack tooling for our Safety Systems teams that directly improves the safety and reliability of OpenAI’s models, including in sensitive areas like mental health and other vulnerable-user protections. Your work will increase the team’s velocity in identifying and fixing safety issues and help tighten the feedback loop between policy, data, and the model training cycle. In this role, you will: Own the end-to-end development of internal tools that help improve the safety of OpenAI’s models (with a focus on areas like mental health and other vulnerable-user protections) Partner closely with Safety Systems researchers, engineers, and model policy creators to understand workflows, pain points, and requirements—and translate them into durable product solutions Build full-stack experiences to support core model policy workflows, such as labeling and inspecting data, analyzing and reviewing failure cases, and surfacing insights for iteration Optimize internal applications f

JavaScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. In this role you will: As a Hardware Test Engineer, you will work on Machine Learning/AI hardware system projects to craft the solutions for current and future data center deployments. You will bring a strong understanding of hardware system testing, excellent project management skills, and the ability to collaborate across multiple teams to ensure efficient lab operations. You will be responsible for designing, implementing, and executing comprehensive test plans that ensure the reliability, performance, and scalability of our supercomputing hardware systems. You will develop detailed test plans and methodologies tailored to hardware components, including processors, memory modules, custom accelerators and interconnects. You will collaborate with hardware design, manufacturing, firmware teams and vendors to identify, analyze, and resolve issues affecting hardware, power, thermal and high-speed interconnects. You will perform in-depth debugging on the hardware system Excellent analytical skills to diagnose hardware issues, troubleshoot problems, and propose solutions. Ability to interpret complex test data, identify trends, and draw meaningful conclusions. High-speed links, with a focus on SerDes (Serializer/Deserializer) technology to assess signal integrity, error rates, and overall link performance. You will collaborate with the lab manager to maintain the equipment and hardware systems, including oscilloscopes, thermal test chambers, liquid cooling systems, and other mea

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo

AWSLinuxRestAI
🔔

Get new reliability engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime