About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high
Jobs in United States
Reliability Engineer in San Francisco
227 active opportunities · Updated October 2026
Showing
15 jobs
Explore current reliability engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Learn more about OpenAI’s approach to safety. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. About the Role At OpenAI, we're dedicated to advancing artificial intelligence, and we know that creating a secure and reliable platform is vital to our mission. That's why we're seeking a software engineer to help us build out our trust and safety capabilities. In this role, you'll work with our entire engineering team to design and implement systems that detect and prevent abuse, promote user safety, and reduce risk across our platform. You'll be at the forefront of our efforts to ensure that the immense potential of AI is harnessed in a responsible and sustainable manner. Your Responsibilities: Architect, build, and maintain anti-abuse and content moderation infrastructure designed to protect us and end users from unwanted behavior. Work closely with our other engineers and researchers to utilize both industry standard and novel AI techniques to measure, monitor and improve AI models’ alignment to human values. . Diagnose and remediate active incidents on the platform and build new tooling and infrastructure that address the root causes of system failure. You might thrive in this role if: You have built and run production services in a high growth, rapidly scaling environment. You can debug live issues and restore systems quickly. You have worked on content safety, fraud, or abuse, or are motivated and excited to work on present-day (“now-term”) AI safety. You have experience with Python or with modern languages such as C++, Rust, or Go, and are able to quickly ramp up on Py
About the Team The Safety Systems team is dedicated to ensuring the safety, robustness, and reliability of AI models and their deployment in the real world. Building on the many years of our practical alignment work and applied safety efforts, Safety Systems addresses emerging safety issues and develops new fundamental solutions to enable the safe deployment of our most advanced models and future AGI, to make AI that is beneficial and trustworthy. Learn more about OpenAI’s approach to safety About the Role As an Analytics Engineer in Safety Systems, you will play a pivotal role in building a data-centric culture, enhancing decision-making processes, and driving strategic initiatives through analytics. You will partner closely with Engineering, Research, and Data Science to develop and maintain canonical data sources and source-of-truth dashboards that enable both people and AI agents across the organization to derive trustworthy, actionable insights. You will own the consumption layer for safety metrics: defining intuitive, reliable ways for stakeholders across Safety Systems, partner teams, and leadership to understand the safety of our products, answer safety-related questions independently, and inform product decisions and company strategy. Most importantly, you will be a core member of the Safety Systems team, collaborating with researchers and engineers to advance our goals of safe, robust, and reliable AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and maintain canonical datasets that serve as sources of truth for safety metrics. Develop and refine data products such as dashboards, reports, agent-enabled workflows, and machine-readable interfaces that empower stakeholders to extract and analyze data independently. Work closely with stakeholders in Engineering, Research, and Data Science to understand their decision-making n
Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are looking for an experienced Mechanical Engineer with 7+ years of experience in design of IT hardware from chip/package to system levels. You’ll work alongside experts in thermal, mechanical, electrical, software, and systems engineering to support the design, analysis, and validation of mechanical and thermal systems that ensure the reliability, efficiency, and longevity of mission-critical hardware. This position requires strong analytical skills, hands-on testing experience, and the ability to work in a fast-paced, cross-disciplinary environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead mechanical design for AI supercomputer product in the data center application Collaborate with the cross functional team to design and optimize thermal solutions for data center hardware, including chips, power modules, and system-level cooling architectures Collaborate with cross-functional teams to integrate thermal management strategies into hardware design, from concept to mass production Design and validate mechanical systems, including chassis, enclosures, cooling systems, and high-power connections, ensuring alignment with performance and reliability standards. Perform 3D modeling, FEA, tolerance analysis, and prototyping, ensuring manufacturability and a
About the Team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn from deployment and broadly distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. We aim to make our innovative tools globally accessible, transcending geographic, economic, or platform barriers. Our commitment is to facilitate the use of AI to enhance lives, fostered by rigorous insights into how people use our products. About the Role We are seeking Software Engineers (Emerging Talent) to join our Applied Engineering team. You’ll work in a highly iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. We value engineers who are self-starters, care deeply about the end user experience, and take pride in building products to solve customer needs. In this role, you will: Own the development of new customer-facing ChatGPT and OpenAI API features and product experiences end-to-end Talk to users to understand their problems and design solutions to address them Collaborate with a cross-functional team of engineers, researchers, product managers, designers, and operations folks to create cutting-edge products Optimize applications for speed and scale Create a diverse and inclusive culture that makes all feel welcome. Your background looks something like: Bachelor's or Master’s degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 0-1 years of experience in software engineering or a relevant field Proficiency with JavaScript, React, and some backend languages (we use Python) Some experience with relational databases like Postgres/MySQL Interest in AI/ML (direct experience not required) Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadl
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team Customer education helps customers and partners build the practical skills and confidence to use AI and OpenAI products safely and effectively. The team focuses on role- and skill-based learning paths, practical content, and product experiences that accelerate learning in the workplace. It brings together learning and enablement expertise, field insight, product signals, and measurement to improve learner and business outcomes. Together, these experiences will help enterprise users build practical AI skills, apply them with confidence in their work, and demonstrate what they can do. Employers will gain a clearer view of workforce skills and progress, helping them recognize capability, focus development where it matters most, and build confidence in workforce readiness. About the Role We’re looking for a full-stack engineer to define and build a new class of learning experiences. This is an early-stage product area where technical judgment, product sense, and learner empathy are critical. You will be setting a technical vision for how people use AI to learn how to use AI, safely and beneficially. This is a hands-on, 0-1 product engineering role with broad technical and product ownership. You’ll set direction, make foundational decisions, and ship the first versions of experiences that can grow into the default way people learn at work. You will drive full-stack product experiences end to end, from prototype through launch, instrumentation, iteration, and production hardening. The work spans interaction design, frontend implementation, backend APIs and services, learner state, content and runtime integration, telemetry, evaluation, reliability, safety, accessibility, and launch readiness. You’ll work closely with our education, GTM, and engineering teams to translate how people learn into products people want to use. bring role- and skill-based learning paths into the product, designing coaching, feedback, and adaptive support which responds to each lea
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role Are you passionate about ensuring the highest quality for cutting-edge generative AI applications? As a software quality engineer at WRITER, you'll play a critical role in shaping the reliability, performance, and trustworthiness of our AI-powered work orchestration platform. You’ll be at the forefront of defining and implementing rigorous quality strategies for our enterprise-grade LLMs and AI agents, directly impacting how hundreds of global companies unlock transformational value through AI. This is a unique chance to dive deep into the unique challenges of AI quality assurance and make a tangible difference in a rapidly evolving field. This is a hybrid role based out of our London, San Francisco, Seattle, and New York City hubs. You will report directly to the director of engineering. 🦸🏻♀️ What you'll do Define and implement comprehensive quality assurance strategies and test plans for our AI agents and LLM-powered applications, ensuring exceptional prod
About the Team OpenAI’s Network Engineering team within IT and Security advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient network services. We build and operate the connectivity that supports OpenAI’s offices, labs, campuses, cloud environments, people, and devices. By combining strong network fundamentals with security, reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at OpenAI, you will design, operate, and continuously improve the global networks that connect our offices, labs, campuses, PoPs, cloud environments, people, and devices. The role spans strategic platform engineering and responsive production operations: you will shape architecture, standards, roadmaps, lifecycle plans, and automation while supporting incidents, escalations, and time-sensitive delivery. Operational signals will inform what we stabilize, simplify, standardize, or automate next. We work backward from user needs, investigate root causes, own outcomes end-to-end, and move quickly without compromising security. We are looking for a versatile engineer who can make pragmatic reliability and security tradeoffs, communicate clearly, and turn recurring operational work into durable platforms, tooling, and standards. You will partner across IT, Security, AppEng, Research, Applied, workplace teams, carriers, and vendors. In this role, you will: Design, implement, and operate secure, scalable enterprise networks across offices, labs, campuses, PoPs, cloud connectivity, and hybrid environments. Set strategic direction for network services through architecture, standards, roadmaps, lifecycle planning, capacity strategy, and measurable reliability outcomes. Own production operations, including on-call, incident response, escalations, and time-sensitive delivery, while protecting user experience,
About the Team Our Cyber team builds AI systems and products that help trusted defenders understand and respond to cyber threats while improving the safety and reliability of frontier models in security-sensitive settings. The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced cyber capabilities useful to defenders and responsibly managed. We collaborate closely with Safety/Preparedness, Research, Security, Legal, Communications, GTM, and external partners across OpenAI’s broader cyber work. About the Role We’re looking for research and software engineers to join Codex Cyber. You’ll help define and ship security products, work with trusted defenders and customers, shape model training and access patterns, and build research and evaluation systems for assessing cyber capabilities, validating safeguards, and improving training data. This role is hands-on and cross-functional, connecting product launches, model development, safety work, and real-world security use cases. In this role, you will: Help define and execute the technical roadmap for Codex Cyber’s security products, including evaluations, safeguards, trusted-defender workflows, and deployment decisions. Work with trusted defenders, customers, and partner teams to understand cyber use cases, evaluate risk, and turn feedback into product and research priorities. Shape cyber-specific model training and access patterns, including data, evaluations, validation, and deployment criteria. Build and validate systems for measuring cyber capabilities, monitoring misuse risk, and proving safeguards work in practice. Collaborate with Safety/Preparedness, Research, Security, Legal, Communications, Go-to-Market, and external partners on company-wide cyber priorities. Translate frontier cyber research into launch-ready tools, operational playbooks, and durable infrastructure for Codex and security products. You might thrive in this role if you: Enjoy 0 -> 1 envi
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and Engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. This role is based in San Francisco, CA, with two additional locations under consideration: London, UK, and Dublin, Ireland. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation o
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
Other cities to consider
More places hiring for this role
Get new reliability engineer jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime