ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a seasoned Frontend Engineer to craft performant and delightful user experiences across Baseten’s core platform. You’ll own critical parts of our web application stack and collaborate cross-functionally with product, design, and backend teams to launch impactful features that help users deploy and manage AI systems at scale. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Core Product team: Rolling Deployments Model APIs for frontier models Model training built for production inference RESPONSIBILITIES Design, implement, and maintain responsive, accessible, and user-friendly frontend interfaces using React and TypeScript Collaborate closely with product designers to turn complex ideas into elegant, intuitive UIs Optimize application performance and reliability, with a focus on rendering speed and responsiveness Drive major frontend initiatives, including partnering with backend teams to define APIs and test and refine end-to-end flows Establish best practices, and mentor other engineers on frontend technologies Build reusable component libraries and frontend infrastructure that accelerate product development Partner with backend and platform teams to define and refine APIs and end-to-end flows REQUIREMENTS 5+ years of experience building production-grade web applications Deep expertise in React, TypeScript, and modern web development tooling Track record of bu
Jobs in United States
Application Reliability Engineer in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current application reliability engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence. In this role you will Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring You’ll love this job if you Care deeply about correctness, rigor, and measurement in AI systems Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team Thrive in cross-functional environments and enjoy influencing model design through better evaluation Qualifications Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learni
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Streaming Platform team at Sentry is building the next generation of infrastructure that powers our ingestion pipelines and real-time data processing systems. Our platform ingests, processes, and distributes hundreds of thousands of events per second with low latency and high reliability. We are creating a system that makes it easy for Sentry engineers to deploy and run Streaming Applications at scale by simplifying the complexity of Kafka, scaling consumers automatically, and managing state so product teams can focus on building great experiences for developers. As part of this team, you will work on challenges at the intersection of distributed systems, real-time data processing, and developer experience. You will help us create a self-service streaming platform that improves stability, accelerates time to production, and reduces operational overhead. In this role you will Design, build, and operate components of our Streaming Platform, including Kafka, the streaming runtime, high-level APIs, and developer-facing abstractions. Implement resilient, high-throughput stream processing systems that handle unbounded datasets with strong correctness guarantees (delivery, checkpointing, watermarking, and more). Build scalable automation and control plane for Kafka fleet management and improve efficiency. Partner with product engineers to ensure our abstractions enable fast, reliable, and consistent ingestion pipelines. Improve observability, monitoring, and failover for mission-critical real-time systems. You’ll love this job if you You enjoy working on distributed systems at scale and care about reliability and
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. About the role As a Senior Software Engineer, you will architect and scale the core systems that power Sentry's billing infrastructure, ensuring accuracy and reliability at massive scale. You will collaborate on building the next generation of Sentry’s usage tracking pipeline, processing hundreds of billions of events daily with low latency and financial-grade accuracy. You will help design flexible pricing primitives that support everything from per-event usage billing to complex enterprise contracts, enabling product and sales teams to experiment rapidly while maintaining revenue accuracy and reduced time-to-market for new products. You will contribute to technical decisions on data consistency challenges unique to billing—like handling event delays, retroactive pricing changes, and distributed count reconciliation across our infrastructure. You'll love this job if you Want to solve the "easy to explain, hard to build" problems—like ensuring a customer's bill matches their usage perfectly, even when processing hundreds of billions of events daily across distributed
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are looking for an experienced Mechanical Engineer with 7+ years of experience in design of IT hardware from chip/package to system levels. You’ll work alongside experts in thermal, mechanical, electrical, software, and systems engineering to support the design, analysis, and validation of mechanical and thermal systems that ensure the reliability, efficiency, and longevity of mission-critical hardware. This position requires strong analytical skills, hands-on testing experience, and the ability to work in a fast-paced, cross-disciplinary environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead mechanical design for AI supercomputer product in the data center application Collaborate with the cross functional team to design and optimize thermal solutions for data center hardware, including chips, power modules, and system-level cooling architectures Collaborate with cross-functional teams to integrate thermal management strategies into hardware design, from concept to mass production Design and validate mechanical systems, including chassis, enclosures, cooling systems, and high-power connections, ensuring alignment with performance and reliability standards. Perform 3D modeling, FEA, tolerance analysis, and prototyping, ensuring manufacturability and a
About the Team The Cybersecurity Products team builds products at the frontier of AI and cybersecurity. Our work includes Codex Security and related cyber products that turn advances in model capability into dependable tools for defenders. We help teams find, validate, and remediate vulnerabilities, continuously improve the security of software, and test AI-powered applications before they reach production. About the Role As a Full Stack Software Engineer, you will build the product experiences and systems that make AI-powered security useful in real engineering environments. You will work across web surfaces, APIs, orchestration, data models, and integrations to help security and engineering teams move from a codebase or application to evidence-backed findings, prioritized remediation, and revalidation. You will collaborate closely with product engineers, security researchers, and customer-facing teams. The work spans fast-moving product development and hard systems problems: long-running workflows, large repositories, sensitive data, reliability, observability, and a high bar for earning user trust. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build end-to-end workflows for vulnerability discovery, security scanning, red teaming, findings review, remediation, and reruns. Design and operate backend services for long-running security work, including APIs, asynchronous orchestration, durable state, and integrations with developer workflows. Make complex security results actionable through clear product surfaces, strong evidence, thoughtful prioritization, and reliable reporting. Partner with security researchers, product teams, and users to evaluate quality, reduce noise, improve coverage, and ship safely. You might thrive in this role if you: Have experience shipping production full-stack products across modern web frontends and backend s
About the Team OpenAI’s Client Platform Engineering (CPE) team delivers trusted devices at scale: secure by default, reliable by design, and effortless to use. We own platform capabilities across macOS, Windows, iOS, Android, and Linux, spanning endpoint posture and device trust, application delivery, onboarding, updates, telemetry, workflow orchestration, and employee-facing remediation. The team partners deeply with Security, Research, Applied, and specialized engineering groups to enable and protect OpenAI while reducing friction for the people advancing our mission. About the Role As an Engineering Manager for CPE, you will lead a team of engineers responsible for the strategy, delivery, and operation of OpenAI’s cross-platform client foundation. You will combine people leadership with strong technical judgment: setting direction, developing engineers, reviewing architecture and tradeoffs, and creating the operating mechanisms that turn ambiguous needs into durable platform outcomes. This is a high-leverage role at the intersection of security, reliability, developer velocity, and employee experience. CPE is a highly technical platform engineering organization delivering first-party services, automation, observability, and safe fleet operations. We’re looking for a leader who can guide its next chapter, scaling the team and its systems, partnering across the company, and raising the bar for secure, reliable, low-friction experiences across every supported platform. In this role, you will: Lead and develop a high-performing engineering team; hire thoughtfully, coach engineers, create clarity, and foster an inclusive, high-accountability culture that pushes perceived limits. Define and execute a multi-year client-platform strategy and roadmap across macOS, Windows, iOS, Android, and Linux, including how Codex and agents can reshape employee computing. Provide technical direction for endpoint posture, device trust, application delivery, device onboarding, updates,
$220K – $450K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in the company. They build the internal control platforms, configuration systems, traffic routing, and automation that let product engineers operate services safely at scale without needing deep infrastructure expertise themselves. As the Engineering Manager for Infrastructure Engineering, you'll lead a team of engineers building the tools that power Sentry's growth: internal admin and change management tools, configuration automation, and the routing layer that underlies Sentry's architecture. You'll be responsible for technical vision, team health, system reliability, and partnership with engineering teams across the company who depend on your team's tools every day. You'll work closely with leaders across Infrastructure, Platform, and Production Engineering to shape how Sentry scales its operational model as the company grows. In this role you will Lead a team of engineers building the internal control platforms that every engineering team at Sentry relies on to operate services safely. Drive the evolution of Infrastructure Engineering's platform, including configuration management, traffic routing and environment controls Own the team's technical direction, contributing to key decisions on API architecture, internal tooling design, and automation frameworks. Nurture and grow engineers at different levels, providing support through coaching, mentorship, and career development. Foster an inclusive, high-performing team culture focused on ownership, learning, and delivery. Partne
About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are looking for a senior, hands-on operator to own Order Management and Billing execution across OpenAI’s cloud marketplace and partner ecosystem, including platforms such as AWS, GCP, Oracle Cloud, GovCloud, and future channels. This senior individual contributor role will translate partner requirements into scalable workflows and ensure launch readiness, accurate billing, and reliable daily execution. As a senior individual contributor within the Cloud Marketplaces team, you will own the end-to-end order-to-invoice lifecycle for your assigned portfolio. You will ensure that private offers, commercial terms, provisioning, pricing, usage, billing data, credits, settlements, and partner-specific reporting flow through our systems accurately, on time, and with audit-ready controls. You will implement and continuously improve the common cloud marketplace operating model, lead cross-functional execution for your assigned portfolio, and surface risks, requirements, and improvement opportunities. You will partner across Revenue Systems, Product, Engineering, GTM, Finance, Partner Operations, and external marketplace stakeholders. This role is critical to building the operational backbone for OpenAI’s expansion across cloud marketplaces and government-cloud channels. You will combine deep operational judgment with process and control execution, automation, clear communication, and hands-on problem solving to improve billing reliability, partner experience, customer outcomes, and financial integrity at scale. This role
About the Team The SaaS and Software Governance team sits within Corporate IT and helps OpenAI scale from startup-speed tooling to mature enterprise architecture. The team owns practical governance for software, SaaS, integrations, APIs, identity, data access, and agent-enabled workflows, with a mandate to improve security, reduce software sprawl, and help business teams move faster through better foundations. About the Team OpenAI is scaling from an emerging, high-velocity startup into a mature enterprise operating model. The IT Software Architect will help shape the software, SaaS, Data integration, and agent-enabled architecture that lets the company move quickly while improving security, compliance, data quality, and customer, partner, and employee experience. This role is not a traditional ivory-tower architecture function. It is a hands-on governance and enablement role that partners with business teams, IT operations, Security, Procurement, Business Platforms, Applied teams and Data Engineering to guide software decisions, reduce unmanaged sprawl, and build reusable enterprise foundations. Why This Role Matters OpenAI’s software footprint is expanding rapidly across SaaS, internally built tools, agents, integrations, APIs, third-party platforms, and application systems that OpenAI. The company needs a stronger tools architecture layer that can help teams make good decisions early, avoid duplicate tools, govern sensitive data and identities, and identify where OpenAI should build instead of buy. The person in this role will help turn software governance from an approval checkpoint into an enterprise capability: a system that improves speed, reliability, security, and business outcomes. What You'll Do Own the target architecture for enterprise software, SaaS, integrations, APIs, and agent-enabled business systems across Corporate IT. Drive deprecation and consolidate targets for enterprise software. Build lightweight governance patterns that guide teams before
About the Team OpenAI’s Platform and Infrastructure Engineering organization advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient technology solutions. Our team builds and maintains robust infrastructure that safeguards OpenAI’s data and systems while ensuring employees are well-equipped and seamlessly connected. By prioritizing security, reliability, and user-centric solutions, we empower OpenAI employees to drive impactful AI research, corporate operations, and product innovation. About the Role As a Software Engineer: Internal Applications, Enterprise, you will build internal products that make technology support and administration safer, faster, and less dependent on manual intervention. You will help reduce reliance on broadly privileged human actions, turn recurring technology problems into paved paths, and build agentic systems that can help resolve tickets end to end. A core part of the role is building the interfaces that bring employees, AI agents, and human responders together in a shared ITSM experience, with the right context, controls, and handoffs at each step. We are seeking engineers who enjoy working across frontend and backend layers on ambiguous, high-leverage enterprise problems. You should bring strong product judgment, solid backend engineering fundamentals, and an interest in building software that changes how technology support, system administration, and agent-assisted operations are delivered. The best fit will care as much about the quality of the operator and employee experience as the correctness of the backend systems behind it. In this role, you will: Build frontend experiences that let employees request help, let agents gather context and take safe actions, and let human responders review, approve, or take over without losing the thread. Reduce reliance on broadly privileged manual actions by replacing them with narrow, auditable, policy-aware aut
Other cities to consider
More places hiring for this role
Get new application reliability engineer jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime