Jobs in United States

Performance And Systems Engineer in San Francisco

364 active opportunities · Updated October 2026

Explore current performance and systems engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan

SQLAWSRestAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Cloud Platform Engineer, you'll envision and build robust systems and processes that ensure our infrastructure is scalable, reliable, and efficient. This can range from automating deployments and monitoring systems to optimizing performance and managing incidents. We all work closely with our users, learning from their past struggles in operationalizing ML, onboarding them onto our platform, and turning our learnings into ideas for improving Baseten. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Build and maintain scalable infrastructure to support the deployment and operation of machine learning models. Establish standards and best practices for reliability and performance across the infrastructure. Automate processes when relevant, particularly for managing CI/CD pipelines. Own products and projects end-to-end, functioning as both an engineer and a project manager, with a focus on user empathy, project specification, and end-to-end execution. Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions. Mentor junior team members and contribute to knowledge sharing within the organization. Navigate ambiguity and exercise good judgment on tradeoffs and

KubernetesCI/CDGitMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About Role We are looking for an Operations Program Manager - Robotics Data Acquisition to own the day-to-day operating rhythm in our data collection facilities. You will work closely with operators, technicians, program managers, and engineers to keep rigs ready, campaigns moving, issues resolved, and performance improving. This is a hands-on operations role that requires you to be comfortable spending time on the floor, working through ambiguity, and using data to make the operation more reliable and efficient. This role is based in San Francisco, CA and requires in-person presence 5 days a week. In this role you will: Coordinate daily operations readiness across workstations, operators, materials. Track core operating metrics including utilization, cycle time, throughput, downtime, operator productivity, and data quality. Identify bottlenecks through workflow analysis, time studies, and capacity modeling, then drive practical fixes. Execute the rollout of new hardware, sensors, tools, and process changes with Engineering, Operations, Facilities, Supply Chain, and Safety. Identify equipment readiness issues and coordinate with technical support to keep workstations, and test equipment calibrated, configured, maintained, and ready for rollouts and evaluations. Lead root cause analysis for recurring operational issues and follow through on corrective actions. Provide operation input to create and maintain SOPs, work instructions, training materials, and process controls. Identify and flag resource constraints and manage issue escala

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Infrastructure organization builds and evaluates the systems that power advanced AI workloads. We work closely with hardware, modeling, and architecture teams to ensure that new platforms deliver real-world performance aligned with workload needs. Our team focuses on understanding workload behavior across evolving hardware platforms—bridging the gap between theoretical capability and observed system performance. About the Role We are seeking a Workload Porting & Performance Engineer to evaluate new hardware platforms by porting benchmarks and real-world workloads, analyzing performance, and identifying system bottlenecks. In this role, you will bring up workloads on new systems, characterize performance behavior, and adapt workloads to better utilize hardware capabilities. You will play a critical role in validating new platforms and ensuring that performance aligns with expectations across compute, memory, and networking subsystems. This role requires strong hands-on experience with performance analysis, workload optimization, and system-level debugging across hardware and software boundaries. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Port and enable benchmarks and real-world workloads on new hardware platforms. Evaluate system performance across compute, memory, storage, and networking subsystems. Identify and analyze performance bottlenecks and inefficiencies. Adapt and optimize workloads to better utilize hardware capabilities. Develop and run performance experiments and profiling workflows. Compare expected vs. observed performance and provide feedback to: hardware architecture teams performance modeling teams system and software engineers. Debug issues across the stack, including software, runtime, and hardware interactions. Provide actionable insights to guide platform readiness and deployment decisions. Qualifications E

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec

TypeScriptPythonJavaSQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work

PythonAWSCI/CDRest
S
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.4%
Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role A strong and reliable platform is essential to scaling Sentry for the future. Our Platform organization is responsible for everything that powers Sentry—from cloud infrastructure and streaming systems to storage, deployment, and security. We own the core services and technical foundations that enable every product and engineering team at Sentry to move fast and build with confidence. We're looking for a passionate and pragmatic Senior Staff Software Engineer to help lead this evolution. In this role, you’ll report directly to the VP of Engineering and collaborate with teams across the company to shape the future of Sentry’s platform. What You’ll Do Architect the future of Sentry by translating business needs and product strategy into clear, scalable technical blueprints. Partner with product and engineering leaders to align technical roadmaps with company goals. Lead cross-cutting initiatives across the Platform org—owning them end-to-end and driving meaningful outcomes. Promote engineering excellence by mentoring platform engineers, sharing best practices, and setting high standards for system design, scalability, and operational quality. Review major architectural proposals and help ensure consistency, maintainability, and long-term technical health across the company. You’ll Love This Job If You... Enjoy designing and building platforms that help teams move faster and scale safely. Thrive on solving complex, multi-dimensional problems across product, infrastructure, and organizational layers. Want to make architectural decisions that shape Sentry’s long-term success. Bring new ideas, tools, and frameworks t

AWSGCPKubernetesAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal

AWSRestAIRust
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Staff Machine Learning Engineer on Sentry’s AI/ML team, you’ll be directly responsible for developing the models and agents used to make our product smarter and more capable. This role is crucial; you will be at the forefront of integrating AI and machine learning into our core products, from issue triage and resolution to predictive analytics for application performance monitoring. Your work will help companies around the globe gain actionable insights into their software, enabling them to build better products, faster. In this role you will Build state-of-the-art agentic AI systems to triage, debug, and solve real production issues Leverage Sentry’s novel (and massive) dataset of errors, spans, and profiles Own the development of major initiatives in the AI/ML space You'll love this job if you Are driven by impact and enjoy working on high-stakes, high-visibility projects Enjoy building things. You will have the opportunity to join the AI/ML team as one of its foundational members Thrive in cross-functional teams and enjoy building features alongside developers and product teams Qualifications Minimum 4+ years of professional experience with a MS/PhD degree in computer science, machine learning, or a related field Minimum 6+ years of professional experience with Bachelor’s degree in computer science, machine learning, or a related field Demonstrated expertise building production-grade agentic systems and tools You are comfortable writing production quality code (we use Python) Expertise with deep learning frameworks (we use PyTorch) Familiarity in deploying machine learning models at scale in production

PythonMachine LearningAIRust
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$220K – $450K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role AI and machine learning are reshaping how developers debug, monitor, and ship software, and Sentry is uniquely positioned to lead that shift. We sit on a novel and massive dataset of real production errors, spans, and logs from tens of thousands of engineering organizations — the kind of signal that makes ML genuinely useful, whether it's a clustering model that groups related issues, a ranking system that surfaces the right alert at the right time, or an agent that proposes a fix. We're looking for an Engineering Manager to lead and grow our Machine Learning Engineering team. This team owns the full spectrum of ML at Sentry: classical techniques like clustering, ranking, anomaly detection, and embeddings that quietly power core product surfaces today, alongside the LLM-based and agentic systems shaping where the product is headed. You'll partner closely with product, design, and engineering leaders to decide where ML belongs in our products, what kind of ML actually fits the problem, and how we translate that work into experiences millions of developers rely on every day. In this role you will Set technical direction across the team's full ML surface area — from classical models for clustering, ranking, and anomaly detection to LLM-based and agentic systems — and make sharp calls about which approach fits each problem Define how the team evaluates and monitors ML systems in production, from offline metrics to online experimentation to model and agent observability Stay hands-on enough to review code and model designs, contribute to architecture discussions, and unblock engineers on complex ML problems Define

RestMachine LearningAIGo
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin

PythonDockerCI/CDRest
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role This isn’t a typical engineering role. You won’t be embedded in a single product team or siloed in one product area. Instead, you’ll sit within Platform Engineering, own the AI-assisted coding domain, and work across all of engineering at Sentry, focused specifically on how AI coding agents participate in our software development lifecycle. For AI coding agents to work well in our repo, the internal systems they depend on need to be accessible via API, not locked behind UIs that require human interaction. Right now, many of those systems aren’t agent-ready. You’ll audit and prioritize that gap, expose those systems programmatically, and build the connections that let tools like Claude Code operate on them end-to-end. From there, the scope expands to improving the quality of AI-generated pull requests and automating the engineering work that’s important but consistently deprioritized. You will look from context engineering standpoint to see what to send to our model; you will look from harness engineering standpoint to see the tools it can use, the permissions it has, the state it carries forward, the tests it has to pass, the logs you capture, the retries, checkpoints, guardrails, and evals. You’ll work closely with the dev infrastructure team as your home base, then collaborate across every product team coding in our repo once the tooling foundation is in place. It’s a broad role with real impact, and the work you do will directly change how Sentry engineers ship software. What You’ll Do Audit Sentry’s internal developer systems and make them API-ready for AI agents. You’ll prioritize and drive the work of ex

CI/CDMachine LearningAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction

RedisAWSRestAI
🔔

Get new performance and systems engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime