Jobs in United States

Reliability Engineer in San Francisco

227 active opportunities · Updated October 2026

Explore current reliability engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of this API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch. This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows. You should be excited to work across research, engineering, data, evals, and product to make models better at acting in real workflows. You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/align

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Operations Program Manager (OPM) to serve as the single-threaded operational leader for new hardware introductions (NPI) and production ramps across OpenAI’s AI infrastructure systems. This role combines hands-on execution with strategic ownership. You will be responsible for defining the operating model, aligning cross-functional stakeholders, setting the critical path, making informed tradeoffs, escalating decisively, and ensuring hardware programs deliver on schedule, quality, cost, and scalability. Success in this role requires comfort operating in ambiguity, influencing without authority, and driving alignment across internal teams and external partners—while keeping eyes firmly on long-term system scalability and repeatability. In this role, you will: Strategic & Leadership Ownership Act as the single-threaded owner for operational readiness across NPI and ramp, accountable for outcomes from early bring-up through sustained production Translate OpenAI’s infrastructure strategy and engineering objectives into clear operating plans, execution priorities, and decision frameworks Drive alignment across Engineering, Operations, Strategic Sourcing, Finance, Capacity Planning, and Executive stakeholders by framing tradeoffs, risks, and recommendations Proactively identify inflection points where decisions or investments are required to protect long-term scale, reliability, or cost targets Influence operational strategy with manufacturing par

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Member of Technical Staff on AI Infrastructure, you will build and maintain the foundational systems and distributed infrastructure that power AI model post training, inference, and data pipelines. You will collaborate with engineering and research teams to ensure performance, scalability, and reliability of critical AI systems. What You’ll Do Design and implement large-scale, distributed AI infrastructure and services Optimize performance for GPU/xPU accelerators and cloud environments Build tools for observability, reliability, and scaling of AI workloads Partner with cross-functional teams to define AI infrastructure requirements and roadmap Contribute to architectural design and system longevity About You Have experience with GenAI infrastructure systems, distributed systems, cloud computing, and high-performance infrastructure Are proficient in programming languages like Python, Go, or similar Understand scaling challenges specific to AI workloads and accelerators Thrive in fast-paced, collaborative engineering environments The reasonably estimated base salary for this role ranges from $256,000.00 to $276,

PythonAIGoRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires effective engagement with public policy stakeholders and the broader community impacted by AI. Accordingly, our Global Affairs team builds authentic, collaborative relationships with public officials and the broader AI policymaking community to inform and support our shared work in these domains. We ensure that insights from policymakers inform our work and – in collaboration with our colleagues and external stakeholders – seek to shape policy so that it aligns with and supports our mission. About the role OpenAI is looking for an Applied Risk Standards Specialist to lead standards work for applied risk in applications, including mental health, youth safety, age assurance, functional efficacy, reliability, and privacy-related risks. You will help turn internal research, safety policies, and evaluation methods into technically sound, measurable, and adaptable standards that build public trust and support responsible deployment. This role sits within our AI standards and global assurance function in Global Affairs. Working closely with Safety Systems, Research, Product, Model Policy, Legal, and third-party evaluators, you will author proposals, negotiate requirements, and represent OpenAI in priority standards bodies. You will take on the drafting and external engagement needed to advance this work, enabling technical experts to focus on the underlying methods and evidence. You will also bring emerging requirements back into internal planning before they become established expectations. Alongside this primary portfolio, you will support the AI Standards and Global Assurance Lead on frontier AI assurance, contributing to risk-management standards, evaluation and benchmarking requirements, and approaches to independent assessment. The role connects technical practice with external standards. It does not replace t

Artificial IntelligenceAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT infrastructure team is responsible for ensuring that our products can serve rapidly growing demand with the performance, reliability, and quality our users expect. This work sits at the intersection of product demand, model deployment, inference, research, fleet, and capacity. The team translates changing product and model needs into clear capacity decisions and safe, scalable launches. About the Role We are seeking a Technical Program Manager to lead the operating system for Chat capacity and model deployment. You will connect demand forecasting and capacity allocation with model readiness, rollout planning, launch coordination, and post-deployment learning. You will also own mode deployment beyond capacity by working with cross functional teams across research, post-training, inference and product to own mainline model deployment. You will bring structure to constrained-capacity decisions, improve the tooling and mechanisms teams use to prioritize demand, and help new models reach users safely and efficiently. Success requires technical depth, sound judgment under ambiguity, and crisp execution across product, research, infrastructure, and operations teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own cross-functional programs for Chat capacity forecasting, allocation, headroom planning, and constrained-capacity operations. Build durable intake, prioritization, and decision mechanisms that connect product demand and model requirements to available serving capacity. Partner

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

$342K – $445K/yr

Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Employee Tech & Experience (ETX) helps people at OpenAI do their most ambitious work. Across Helpdesk, Executive Support, Systems Operations, Logistics and AV, we make technology simple, reliable and secure. Employee needs guide what we build, improve and choose to eliminate. About the Role Reporting to the Head of Global IT, you’ll lead ETX globally, building on the team’s capabilities and customer-zero work to continually advance the employee experience. You’ll shape ETX’s strategy, investment priorities and operating model in partnership with leadership across the company, turning new capabilities into measurable amplification. You’ll develop leaders and strengthen teams where people feel valued, own meaningful work and enjoy working together. This role is based at our San Francisco headquarters and requires an in-office presence. In this role, you will: Lead the next stage of ETX’s global growth across Helpdesk, SysOps, Logistics and AV, with a shared strategy and accountability for employee outcomes. Continually elevate the employee experience through research and design, directing investment to simplify entire user journeys, remove unnecessary effort and amplify what employees can accomplish. Accelerate ETX’s agent-led and customer-zero work as capabilities advance: continually challenge which workflows need to exist, extend what agents can own end to end and evolve the operating model to deliver measurable amplification. Develop leaders who earn trust, grow others and sustain an environment where people feel valued, take pride in their work and enjoy working together. Give people meaningful ownership and opportunities to stretch and grow, with clear priorities and sustainable workloads. Scale global services to support company growth, with clear commitments to reliability, security, responsiveness, and effective controls. Measure performance through employee effort, service quality, and time to resolution. Shape priorities and investment wi

Artificial IntelligenceAILogistics
P
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About the Team Our Network Enablement and Access team works to unlock the potential of Plaid's network by broadening and deepening our connections with data partners. We build the capabilities that help data providers participate in the network, strengthen the quality and reliability of those connections, and enable great products and experiences for Plaid's customers. Within Network Enablement and Access, the Data Supply Traffic and Health team owns how Plaid's requests flow to the data providers we depend on. We manage the load placed on each provider, the constraints that shape our access, and the fair allocation of capacity across Plaid's products and new initiatives. As paid access expands across the network, we also work to keep that traffic reliable, efficient, and cost-effective. As the Product Manager for Data Supply Health and Traffic, you will establish and lead a new product area at the foundation of every Plaid product. You will define how Plaid allocates constrained provider capacity, scales traffic across products, and manages the economics of paid data access. You will also optimize for data freshness, balancing timeliness with provider capacity and cost so Plaid's

O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. The Power Execution team owns the strategy and execution required to secure reliable, scalable, and economically resilient power for OpenAI’s global data center portfolio. The team sits at the intersection of commercial, technical, policy, legal, and operational work, partnering across OpenAI and with utilities, grid operators, regulators, counterparties, and public-sector stakeholders. About the Role The Energy Regulatory Lead will own energy regulatory strategy and execution for OpenAI’s infrastructure growth. This role will be the primary bridge between the Power Execution team and Public Policy and Government Affairs on energy regulatory matters, ensuring that OpenAI’s external engagement is grounded in project realities and that changing policy and regulatory conditions are translated into actionable infrastructure decisions. This is an individual contributor lead role and does not have direct reports initially. The role combines portfolio-level regulatory positioning with transactional regulatory work: evaluating jurisdictional pathways, supporting utility and energy transactions, coordinating approvals and filings, and helping project teams navigate tariffs, interconnection, load-service requirements, market rules, and regulatory risk from diligence through execution. In this role, you will: Develop and maintain OpenAI’s energy regulatory strategy across priority U.S. markets and, as needed, emerging geographies for infrastructure expansion. Coordinate closely with Public Policy and Government Affairs to shape energy regulatory priorities, engagement plans, messaging, and positions before utilities, public utility commissions, grid operators, state energy offices, and other relevant policymakers. Translate project requirements—load size, timing, reliability, cost, carbon, and expansion needs—into clear regulatory objectiv

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make rigorous, evidence-based talent decisions and ensures that the systems supporting those decisions are valid, reliable, fair, and accountable. About the Role As a People Data Scientist focused on AI fairness and bias testing, you will help establish how OpenAI evaluates AI-assisted People systems and high-impact talent processes. You will design and conduct rigorous assessments to identify, measure, and mitigate potential bias across the lifecycle of models, agents, decision-support tools, and automated workflows. Your work will span the entire employee life-cycle, such as hiring, performance, promotion, employee development, workforce planning, etc. You will evaluate both technical systems and the broader human-AI decision processes in which they operate, examining not only model performance but also data quality, measurement validity, differential outcomes, human oversight, and unintended consequences. We’re looking for an experienced data scientist or applied researcher who can translate complex fairness questions into defensible evaluation strategies, scalable testing infrastructure, and clear recommendations for technical teams and senior leaders. This role is preferred to be based in San Francisco, CA. In this role, you will: Define and lead fairness and bias-testing strategies for AI-assisted People processes, models, agents, and decision-support systems from development through deployment and ongoing monitoring. Design rigorous algorithmic audits and validation studies, including adverse-impact analysis, subgroup and intersectional evaluation, error-rate analysis, calibration, measurement invariance, reliability, criterion-related validity, and sensitivity testing. Identify the appropriate fairness criteria for each use case, evaluate tradeoffs among competing definitions

PythonSQLAWSRest
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s Account Verification team builds the foundation of trust for open finance. We help fintechs and financial institutions connect and verify bank accounts securely so that money can move safely and instantly. Account Verification is the entry point for all of Plaid’s payment-focused consumer experiences and is a critical building block for our customers. The team obsesses over creating seamless verification journeys that balance speed, reliability, and security, enabling consumers to confidently connect to the financial ecosystem. As a PM for Account Verification, you’ll own one of Plaid’s most critical and high-impact product areas. You’ll lead the evolution of our verification platform across Auth, Balance, and Identity Match, defining how millions of people and businesses connect their financial accounts every day. We are looking for a high-ownership builder who thrives in ambiguity, loves building with customers, and is excited to define what’s next for one of Plaid’s most established and strategically important product lines. You’ll set vision and strategy, drive execution across a cross-functional team, and shape how Plaid competes in an increasingly complex and, eventually, AI-driven ver

AWSAIGoRust
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Own day-to-day operations for the scanner program and scanner builds in the spa: purchasing/procurement, vendor management, receiving, inventory, and logistics. Stand up lightweight production operations as we move from prototypes to repeatable builds: build planning, kitting, work instructions, and readiness checklists. Partner with Quality to ensure the operational system supports compliance: traceability, document control, training records, NCR/CAPA workflows, and audit readiness. Drive cross-functional execution for the physical spa build-out and scanner integration: schedules, dependencies, risk register, and weekly coordination with vendors and internal teams. Own the “integration glue” across facilities + device ops: commissioning plans, acceptance criteria, and operational handoff (runbooks, maintenance, spares, escalation paths). Build and track operational metrics: cost, budget, lead times, vendor performance, build throughput, and reliability of critical subsystems. Audit of import/export documentation, management of contract renewals, regulatory compliance. What we’re looking for Proven operations leadership in hardware/medical/robotics (or similarly complex electromechanical products), including procurement and vendor management. Strong program management instincts: can run schedules, unblock cross-functional dependencies, and keep priorities clear under ambiguity. Comfort operating in quality/regulatory environments and building processes that are rigorous without slowing a small team. High ownership and bias to action: can jump between spreadsheets, docks, and the lab/site to keep the program moving. Useful experience Experience running prototype-to-production transitions (NPI, EVT/DVT/PVT-style builds, CM/EMS collaboration). Facilities / construction operations experience (GC coordination, MEP commissioning, site readiness). Familiarity with inventory systems and procurement tooling (even “simple but disciplined”).

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s People Experience & Technology (PXT) team owns the core people platform that powers worker, recruiting, contingent, approvals, and lifecycle workflows across the company. PXT is responsible for operating Workday, Ashby, and related people systems as governed, reliable sources of truth, while building the controls, monitoring, documentation, and auditability required to support scale. About the Role We’re hiring an Enterprise Systems Manager, Recruiting Systems to help own and harden OpenAI’s recruiting platform, with a focus on Ashby and its connected workflows. This is a hands-on systems role for someone who can translate recruiting process problems into governed, durable fixes through configuration, workflow design, access controls, documentation, reporting guardrails, and integration partnership. You will work at the boundary of Recruiting, HR Operations, Legal, Compensation, Analytics, IT, and PXT to improve the reliability and control health of recruiting workflows. The right person is comfortable going deep in system design while also driving rollout, adoption, and operational clarity. In this role you will: Own specific recruiting workflow domains in Ashby and adjacent tools, including stages, fields, permissions, approvals, templates, and configuration standards. Partner on high-priority remediation work across start dates, offers, approvals, integrations, auditability, data integrity, and workflow controls. Design and implement governed workflow changes that balance recruiter usability with reporting trust, downstream integration reliability, and control requirements. Establish and maintain guardrails such as required and conditional fields, stage definitions, role-based permissions, approval logic, validation patterns, and change standards. Drive durable fixes for recurring operational issues by identifying root causes and resolving them through configuration, automation, documentation, or process redesign. Partner with PXT, IT,

AWSRestAIGo
🔔

Get new reliability engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime