Jobs in United States

Performance And Systems Engineer in San Francisco

364 active opportunities · Updated October 2026

Explore current performance and systems engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Human Data Team creates custom data solutions driving groundbreaking research. Our work enhances and evaluates our flagship models and products like ChatGPT, GPT-5, and Sora, and contributes to safety initiatives through collaboration with our Preparedness and Safety Systems teams. About the Role As a Research Program Manager (RPM) in the Human Data team you will partner with research and engineering to design and implement pragmatic solutions for collecting high-quality data. You will be a key interface between our research roadmap, external vendors, AI trainers, and the Human Data engineering team. This role is based in our San Francisco HQ. In this role, you will: Collaborate with Research: Partner with researchers to scope data collection needs, define success metrics, and establish quality measurement frameworks. Design & Execute Data Collection Campaigns: Translate research needs into actionable plans and accelerate execution by leveraging existing tooling and iterating to reach the desired outcome. In many cases, you will need to implement scrappy new solutions while partnering with engineering to design robust/scalable solutions. Unblock Yourself: You must be deeply uncomfortable with the idea of sitting around waiting for external dependencies, and have the technical acumen and drive to figure out how to achieve at least partial success in the interim. Optimize Systems & Processes: Build and optimize dashboards to track campaign performance, leveraging SQL and Python for data analysis and actionable insights. Drive Technical Roadmaps: Collaborate with engineers to enhance data platforms, resolve blockers, and ensure security best practices such as access management. Scale Your Impact : Advise and empower program managers and vendors to drive day-to-day execution so that you can focus on addressing high priority opportunities. You might thrive in this role if you: Are proficient in SQL and Python for data analysis, including q

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Future of Computing Research team is an applied research team within the Consumer Devices group focused on developing new methods, models, and evaluation frameworks that support our vision for the future of computing. We work at the frontier of multimodal AI, helping turn emerging model capabilities into product experiences that are useful, delightful, and worthy of long-term trust. Our work explores a new class of AI systems that can learn over time, adapt to individuals, and support people in the flow of daily life. This includes long-term memory, user modeling, and personalization systems that are aligned not just with immediate satisfaction, but with a person’s broader goals, values, and well-being. We work closely across research, engineering, design, product, and safety to define what it means to build AI systems that know you over time, act at the right moment, and help in ways that are context-aware, respectful, and demonstrably beneficial. About the Role We are looking for a Research Engineer / Scientist to join the Future of Computing Research team to work on RLHF and post-training for personalized, multimodal AI systems. This role will focus on building the learning and evaluation foundations that help models become more context-aware, adaptive, and useful over time. You will work on problems such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that must make high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded: success is not just higher benchmark performance, but better model behavior in real-world use. The ideal candidate is excited about pushing beyond one-turn assistant behavior toward systems that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value. Internally, that maps closely to the need for careful reward design, feedback loops, and evaluation frameworks that test whether i

AWSRestMachine LearningAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Head of IT at Baseten, you will build, scale, and secure our internal technology function to support our rapid growth. Reporting to our Chief Information Security Officer, you will lead and mentor a team of 5+ IT engineers, leading the charge to transition Baseten from startup-era IT to a highly automated, enterprise-ready IT organization. You will take full ownership of corporate IT infrastructure, Helpdesk operations, corporate identity management, device lifecycles, and vendor procurement. As we scale to support the world’s most dynamic AI companies, you will ensure our internal systems scale seamlessly with our headcount, providing a secure, frictionless, and world-class technology experience for all Baseten employees. RESPONSIBILITIES Team Leadership: Manage, mentor, and grow a team of IT engineers, fostering a high-performance culture focused on technical excellence and end-user satisfaction. Helpdesk Operational Excellence: Build a fast-response support function by establishing clear response SLAs, tracking employee satisfaction metrics, and formalizing on-call and incident response processes. Zero-Touch Automation: Architect and implement automated employee onboarding, offboarding, and role-based access changes through deep integrations across HRIS, MDM, and IAM systems. SaaS Management & Procurement: Establish comprehensive SaaS management processes to eliminate shadow IT, automate access w

Machine LearningAIGoRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building for talent density. We believe attracting and retaining exceptional people, and ensuring they feel recognized and valued for their impact, is core to becoming the best place to work. Compensation is a critical lever in that mission. As our Compensation Manager, you will own compensation programs company-wide. You’ll be a trusted advisor to senior leaders, shaping our compensation philosophy, leveling framework, and equity programs to ensure we remain competitive, principled, and performance-oriented as we scale. This role blends strategy and execution: designing clear, fair systems while moving quickly in a high-growth environment. RESPONSIBILITIES Own and evolve Baseten’s company-wide compensation strategy, philosophy, and programs. Collaborate with leadership and HRBP to create and evolve job architecture and leveling frameworks. Build and maintain compensation bands. Conduct regular market benchmarking to ensure comp bands and strategy remain competitive in a fast moving industry. Partner closely with Talent to design and approve competitive new hire offers, advising on negotiation strategy within our compensation principles. Lead bi-annual leveling and compensation review cycles to ensure market competitiveness and reward high performance across teams. Manage new hire equity grants in partnership with Finance and Legal. Design and administer a thoughtful equity refresher program for ten

Machine LearningAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's mission is to ensure that artificial general intelligence benefits all of humanity. The Consumer Devices team is building a new generation of AI-powered products that seamlessly integrate hardware and software to create intuitive, transformative experiences. We bring together experts across embedded systems, machine learning, hardware, design, and product engineering to develop products at the intersection of AI and consumer technology. About the Role OpenAI is seeking a System Performance Engineer to profile, benchmark, and optimize performance across our embedded hardware products. In this role, you will work across operating systems, applications, camera and vision, graphics, and platform teams to define product KPIs, build performance tooling, and drive optimizations from early lab characterization through product launch and real-world usage. You will help establish the performance standards that shape the user experience of our products, ensuring they remain responsive, efficient, and reliable throughout their lifecycle. This role requires deep expertise in embedded or high-performance systems, strong operating systems fundamentals, and hands-on experience debugging under tight latency, power, and memory constraints. This role is based in San Francisco, CA. We use a hybrid work model of four days per week in the office and one day working remotely. Relocation assistance is available for new hires. In this role, you will: Develop system performance benchmarks, methodologies, and policies to evaluate end-to-end product behavior. Profile and analyze performance across key product use cases and workloads using custom and industry-standard profiling tools. Partner closely with engineering teams to identify bottlenecks and drive performance optimizations across the software stack. Define high-level product KPIs and establish measurement frameworks to measure launch readiness and monitor performance throughout the product lifecycle. Measure, re

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a product manufacturing & quality engineer, who will be responsible for driving technical initiatives related to the manufacturing, quality and reliability of our AI supercomputer hardware systems to ensure product success from concept to launch and through mass production. You’ll have the opportunity to coordinate with functional SMEs and work with a wide range of stakeholders, from design engineering and operations teams, TPMs, external industry vendors and partners to ensure that all products are developed and delivered on time and to the highest quality standards. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Own the integrated manufacturing and quality readiness for a product across L6, L10, and L11, with clear gates, milestones, deliverables, owners, and closure criteria. Lead readiness of process flows, tooling, fixtures, assembly operations, test interfaces, and production controls. Review and contribute to work instructions. Translate product requirements into qualification plans, process controls, test requirements and acceptance criteria with design engineering and Area SMEs Coordinate and drive execution of product and process qualification, reliability testing, and validation with the relevant SMEs. Maintain traceable evidence that assigned products and processes meet agreed performance, reliability,

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI develops models that can reason through complex problems and hardware designed for the demands of advanced AI. AI for Chips connects these efforts: applying increasingly capable AI systems to the work of semiconductor engineering. Our goal is to help engineers develop better chips and shorten design cycles. This work brings research, model training, and hardware expertise together to build tools that engineers can use on real designs, with correctness and measurable performance at the center. About the Role We’re hiring a Research Engineer to help OpenAI models solve chip-design problems through reinforcement learning, tool use, and evaluation. You’ll own experiments from the initial idea through implementation and analysis. That means building environments and evaluations, running training, investigating failures, and using the results to decide what to try next. You’ll also build the software needed to make those experiments reliable and reproducible. We value strong coding fundamentals, careful experimental judgment, and the ability to make progress independently. Prior chip-design experience is helpful, but you can learn the domain alongside the team’s hardware specialists. In this role, you will: Build RL environments and evaluations for tasks such as RTL generation, design verification, and physical design optimization. Develop and test approaches that help models use chip-design tools and improve power, performance, and area while preserving correctness. Design experiments, establish baselines, and measure whether improvements hold up on new tasks and designs. Investigate failures across model behavior, rewards, evaluation tools, and experiment infrastructure. Improve iteration speed through better tooling, faster evaluations, and proxy rewards that reflect the outcomes we care about. Turn successful experiments into reusable research code and training workflows, working closely with researchers and engineers. You might thrive in this ro

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the role: We are seeking an experienced Optical Network Engineer to lead Laser related work within our optical interconnect efforts for large-scale compute systems. The role also requires broad, hands-on optical validation experience across IM/DD-based interconnects, working from lab characterization through production readiness and scaled deployment. In this role you will: Drive laser-focused requirements and technical direction within the broader optical interconnect roadmap. Lead evaluation and validation of optical components and subsystems, including laser-based elements, in lab and production-representative environments. Support end-to-end optical testing for IM/DD interconnects (e.g., module/system bring-up, characterization, debug, and readiness for scale). Work with external partners to align on development milestones, performance targets, and quality expectations. Own technical issue triage and resolution across performance, reliability, and manufacturability topics. Collaborate across internal teams to support integration, rollout, and operational success at scale. You might thrive in this role if you have: Strong experience in laser-focused optical engineering (development, validation, manufacturing readiness, or field support). Broad hands-on background with IM/DD optical technologies and optical test/debug workflows. Experience working with external suppliers/manufacturing partners and production-oriented execution. Demonstrated ability to debug complex t

AWSRestAIRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications. You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Core Engineering Responsibilities Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap Performance & Innovation Impl

AWSMachine LearningAIC++
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Infrastructure Software Engineer at Baseten, you'll build and maintain components of our ML inference platform that powers production AI applications. You'll contribute to the core infrastructure, enabling developers to deploy, scale, and monitor ML models with high performance. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Develop infrastructure components for our ML inference platform using Python and Go Implement and maintain Kubernetes deployments for model serving Contribute to our inference orchestration layer for model deployments Build and enhance monitoring systems for model performance metrics Implement efficient resource management solutions for ML workloads Support infrastructure automation to improve ML deployment workflows Work closely with team members to implement technical solutions Help balance performance optimization with system reliability Participate in technical discussions around infrastructure improvements Learn and apply infrastructure best practices REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Go proficiency is a plus Working knowledge of Kubernetes and containeriza

PythonKubernetesRestMachine Learning
N
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -86%

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: The Web Infrastructure team builds the foundations behind Notion’s web clients, including client architecture, performance, reliability, and shared design systems. As Engineering Manager, you’ll lead the team through the evolution from Notion Clients. You’ll set strategy, develop senior engineers and managers, and partner across Notion to make the product faster, more reliable, and easier to build. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: You'll build and manage a diverse and inclusive team of engineers and managers working on core parts of Notion's architecture. You'll create a healthy environment in your team that embodies Notion's values. You'll recruit, coach, and develop engineers. You'll ensure engineers are regularly receiving feedback and making progress on personal and professional goals. You'll facilitate planning—the prioritization, sequencing, and staffing of work—for your team. You'll be responsible

M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Partner with medical image reconstruction scientists / engineers to build ML components that improve reconstruction quality, speed, robustness, or quantitative accuracy. Define training/evaluation pipelines, datasets, and metrics that map to user needs and design requirements. Productionize models: inference performance, reproducibility, monitoring for drift/regressions, and safe fallbacks. Collaborate on hybrid algorithms, incorporating physics and learned priors, denoisers, learned regularizers, and quality estimation. Help build tooling for rapid experimentation as well as rigorous verification of algorithm changes. What we’re looking for Strong applied ML experience plus comfort with signal processing / imaging or adjacent domains. Ability to move fluidly between research prototypes and production-quality systems. Strong evaluation discipline: metrics, ablations, data leakage avoidance, and reproducibility. A demonstrated track record of applying ML to physics-based or inverse problems (i.e., shipped projects, a portfolio, or publications.) Useful experience ML for imaging/inverse problems (or adjacent) with strong evaluation discipline and comfort with GPU performance constraints. Pragmatic production mindset: reproducible training/inference, regression testing, and safe deployment in high-stakes contexts. A background in computational physics or scientific computing. Leverage ML-based methods such as PiNNs and Neural Operators to solve partial differential equations arising in ultrasound simulation and imaging. Experience in Agentic-SciML is a plus. Hands-on experience with data curation for ML: building datasets from messy, real-world sources, defining ground truth, and managing labeling or simulation pipelines. Background in data assimilation: combining observations with physics-based models (Kalman filtering, variational methods, ensemble approaches, or learned variants).

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do. We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research. About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production. This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role, you will: Design and implement inference infrastructure for large-scale multimodal models. Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs. Enable experimental research workflows to transition into reliable production services. Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities. Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers. You might thrive in t

AWSRestAIRust
🔔

Get new performance and systems engineer jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime