About the Team OpenAI is building Procurement of the future: one that uses AI-enabled systems, clean data, and scalable processes to help the business absorb increasing complexity, speed, volume, and throughput while enabling teams across OpenAI to operate more strategically and at greater scale while managing risk responsibly. The Procurement Operations team supports the operational flow of purchasing across OpenAI, helping teams turn business needs into complete, policy-aligned, and executable procurement activity. We partner across Procurement, Finance, Legal, Security, Accounting, Accounts Payable, Systems, and business teams to improve how requests, suppliers, purchase orders, approvals, data, documentation, and downstream handoffs move through our systems. As OpenAI scales, we are building procurement operations that can absorb increasing complexity without adding unnecessary friction. That means using AI thoughtfully, simplifying where we can, strengthening operating discipline where we must, and continuously improving how our stakeholders experience procurement. About the Role We are hiring two Senior Managers, Procurement Operations to manage the operational backbone of how OpenAI turns requests into committed, controlled, and payable spend. One role will focus on Technology categories, and the other will focus on Corporate Services categories. These roles are not removed from the work. You will be expected to review requests in the system, make or guide request and PO updates, resolve operational issues, and use that firsthand view to identify what needs to change. You will push the envelope without losing the thread on risk, staying close enough to the details to understand where friction, ambiguity, and control gaps show up. You will help improve procurement outcomes across intake, purchase requests, PO quality, supplier coordination, approval workflows, documentation, service delivery, and downstream financial handoffs. You will oversee and improve oper
Jobs in United States
Service And Engagement Team Leader in San Francisco
417 active opportunities · Updated October 2026
Showing
15 jobs
Explore current service and engagement team leader jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team The Cloud Agents team builds product infrastructure for long-running agents in the cloud: orchestration, sandboxing and isolation, secure environment connectivity, secrets and identity, observability, reliability, and cost controls. These agents securely connect to diverse developer and customer environments and use tools to accomplish goals. We partner closely with product, research, and infrastructure teams to turn agentic capabilities into dependable platforms for OpenAI products and developers building on OpenAI. About the Role We are looking for an experienced software engineer to help build and scale our cloud agent platform. You will design and operate systems for orchestrating agents at scale. You will work closely with product engineers on ChatGPT, API, and Codex to define the right abstractions and enable them to ship products quickly. Strong backend or infrastructure experience is important; experience with Python, Rust, distributed systems, cloud infrastructure, or product platforms is especially helpful. In this role, you will: Design and scale the orchestration, sandboxing and storage systems that run agentic workloads for Codex, ChatGPT, and the OpenAI API. Partner with product engineers to build a platform that enables them to ship quickly and turn feedback into robust abstractions. Improve reliability, security, performance, and cost efficiency for long-running agents. Deploy services that can operate across different environments and clouds. Your background might look something like: 9+ years of professional engineering experience, excluding internships, in relevant roles at technology and product-driven companies. Experience leading large-scale backend, platform, or infrastructure projects from ambiguous problem statements to production systems. Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, and the ability to move across service, platform, and product boundaries. Strong understanding
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Strategic Sourcing & Procurement function plays a critical role in enabling OpenAI to deliver impact across research, product development, technology infrastructure, and services by helping the company scale responsibly, securely, and with strong commercial discipline. Our work sits at the intersection of innovation and execution. We partner closely with teams across OpenAI to translate rapidly evolving business needs into scalable, compliant, and economically sound external partnerships. As OpenAI continues to grow at pace, services sourcing is becoming increasingly strategic across the company. Every business unit relies on external service providers in different ways — to extend capacity, access specialized expertise, support operations, and accelerate execution. Done well, Procurement becomes a source of trust and momentum, helping OpenAI move faster with the right partners, stronger commercial outcomes, and the right level of protection. About the Role We are seeking an experienced Strategic Sourcing (GTM) Leader to lead strategic sourcing and commercial enablement for OpenAI’s Go-to-Market organization across B2B and B2C channels. You will manage substantial and rapidly growing spend while shaping sourcing strategies and scalable commercial pathways across Media, Creative, Production, Influencer, Agency, Sponsorships, Analytics, Communications, and Event suppliers in support of high-impact global initiatives. You’ll help evolve our GTM procurement function from reactive deal support into a speed-enabling, scalable commercial engine that delivers cost efficiency, launch readiness, and strong governance in a fast-moving environment. In this role, you will: Develop and execute sourcing strategies across GTM, Brand, Global Affairs, Events, Growth, and Partnership activities—spanning both B2B and B2C channels—that align with our mission and b
About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our
About the Team The Cooperative AI team is scaling OpenAI with OpenAI. We are building an AI powered knowledge system that evolves and learns as our products, systems and customers evolve. We leverage our state of the art models, technologies, and products (some external, some still in the lab) to assist or completely automate robust operations supporting both internal and external customers. We support OpenAI customers and internal partners globally, powering systems from customer support to integrity to product insights. We are a self-contained multi-disciplinary team, who enjoy a lightning fast feedback loop with customers at scale, some of whom sit just a few pods away. We iterate fast, and engineer for reliable long-term impact. We're constantly looking for the similarities and patterns in different types of work, and focus on building simple primitives, to apply world class knowledge to many domains. The work of this team exemplifies use of OpenAI technologies. We build systems so everyone can see the leverage that is possible with well designed AI-based implementations. We do this by working through internal use cases focused on Customers (specifically knowledge systems, automation systems, and automated agent systems) to prove impact, then we scale. About the Role We’re looking for Software Engineers who're passionate about blending production-ready platform architecture with new tech and new paradigms. You’ll push the boundaries of OpenAI’s newest technologies to enable interactions and automations that are not only functional, but delightful. We value proactive, customer-centric engineers who can get the foundational details right (data models, architecture, security) in service of enabling great products. In this role, you will: Own the end-to-end development lifecycle for new platform capabilities and integrations with other systems Collaborate closely with engineers, data scientists, information systems architects, and internal customers to understand th
About the Team ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic patterns, complex dependency graphs, and frequent change. The ChatGPT Infrastructure team builds and operates the platforms that enable fast iteration without compromising performance or reliability. We design shared systems, data paths, rollout mechanisms, and reliability guardrails that teams rely on to ship changes to ChatGPT at scale. We focus on high-leverage infrastructure: primitives and “golden paths” that incorporate operational lessons as defaults, so engineers don’t need to rediscover failure modes, latency pitfalls, or integration issues each time they build something new. About the Role We’re hiring Senior and Staff Engineers to design and build infrastructure systems that underlie ChatGPT and multiply the effectiveness of teams building user experiences. This is not a support-only role. It’s a platform-building role: you’ll define interfaces, develop core abstractions, and create tooling to make safe, fast iteration the norm. Your work will reduce friction, prevent regressions, improve performance, and ensure systems scale gracefully as the product grows. Where You Can Have Impact You might work on one or more of the following areas (without being restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution. Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases. Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression contr
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Senior Backend Engineer on the Cloud Platform team, you will play a key role in building the core systems and services that power Postman’s internal platform. You’ll help create new backend services that manage how we deploy, scale, and operate our infrastructure and product services, leveraging Java, Spring Boot, and Hibernate (JPA) on top of cloud-native technologies like Kubernetes, ArgoCD, Istio, and Terraform. This role is highly impactful: the systems you build will be used across Postman engineering, enabling faster delivery, better scalability, and a stronger developer experience. You’ll also have the opportunity to contribute to open source, shaping tools that extend beyond Postman’s boundaries. What You’ll Do Design and develop backend services in Java and Spring Boot to support Postman’s internal Cloud Platform. Architect new services that manage service deployment, lifecycle, and scaling across Kubernetes clusters. Implement GitOps workflows (ArgoCD) to support continuous delivery. Integrate with cloud-native tooling such as Istio, Helm, and Terraform. Apply strong soft
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid is looking for customer-facing professionals to manage and grow Plaid’s relationships with high-growth FinTech customers. The right candidate has the ability to grow a customer’s revenue, build strong partnerships with their customers, develop a deep understanding of their priorities, motives and key drivers, and is solutions oriented. Qualifications Demonstrated relationship management and sales experience. Customer empathy and deep desire to see customers succeed. Desire to get technical and the ability to learn the ins and outs of Plaid's APIs. Some familiarity in financial services and technical products; a high degree of intellectual curiosity. Excitement to work in a high-growth, ever-changing environment and to help build processes and tools as needed. Ability to work with all types of people and build long-lasting relationships. Strong Sales experience (negotiation/value-selling). Executive level communication skills - written and verbal. Relationship building (understand how to leverage internal/external executives, product, etc). Retention and Termination-save tactics. Responsibilities Own and manage the overall relationship with your portfolio of customers in our fastest growing cus
About the Team Customer education helps customers and partners build the practical skills and confidence to use AI and OpenAI products safely and effectively. The team focuses on role- and skill-based learning paths, practical content, and product experiences that accelerate learning in the workplace. It brings together learning and enablement expertise, field insight, product signals, and measurement to improve learner and business outcomes. Together, these experiences will help enterprise users build practical AI skills, apply them with confidence in their work, and demonstrate what they can do. Employers will gain a clearer view of workforce skills and progress, helping them recognize capability, focus development where it matters most, and build confidence in workforce readiness. About the Role We’re looking for a full-stack engineer to define and build a new class of learning experiences. This is an early-stage product area where technical judgment, product sense, and learner empathy are critical. You will be setting a technical vision for how people use AI to learn how to use AI, safely and beneficially. This is a hands-on, 0-1 product engineering role with broad technical and product ownership. You’ll set direction, make foundational decisions, and ship the first versions of experiences that can grow into the default way people learn at work. You will drive full-stack product experiences end to end, from prototype through launch, instrumentation, iteration, and production hardening. The work spans interaction design, frontend implementation, backend APIs and services, learner state, content and runtime integration, telemetry, evaluation, reliability, safety, accessibility, and launch readiness. You’ll work closely with our education, GTM, and engineering teams to translate how people learn into products people want to use. bring role- and skill-based learning paths into the product, designing coaching, feedback, and adaptive support which responds to each lea
About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role A strong and reliable platform is essential to scaling Sentry for the future. Our Platform organization is responsible for everything that powers Sentry—from cloud infrastructure and streaming systems to storage, deployment, and security. We own the core services and technical foundations that enable every product and engineering team at Sentry to move fast and build with confidence. We're looking for a passionate and pragmatic Senior Staff Software Engineer to help lead this evolution. In this role, you’ll report directly to the VP of Engineering and collaborate with teams across the company to shape the future of Sentry’s platform. What You’ll Do Architect the future of Sentry by translating business needs and product strategy into clear, scalable technical blueprints. Partner with product and engineering leaders to align technical roadmaps with company goals. Lead cross-cutting initiatives across the Platform org—owning them end-to-end and driving meaningful outcomes. Promote engineering excellence by mentoring platform engineers, sharing best practices, and setting high standards for system design, scalability, and operational quality. Review major architectural proposals and help ensure consistency, maintainability, and long-term technical health across the company. You’ll Love This Job If You... Enjoy designing and building platforms that help teams move faster and scale safely. Thrive on solving complex, multi-dimensional problems across product, infrastructure, and organizational layers. Want to make architectural decisions that shape Sentry’s long-term success. Bring new ideas, tools, and frameworks t
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin
About the Team The Infrastructure Engineering function sits within IT and is responsible for reliably building, deploying, and operating critical on prem and hybrid environments that power internal services and critical R&D environments. This is an early, high-leverage technical role focused on applying strong Site Reliability Engineering discipline to environments where uptime, safety, recoverability, and security are non-negotiable. This person helps replace bespoke, one-off infrastructure with standardized infrastructure-as-code building blocks that compound reliability and operational leverage as OpenAI scales. About the Role We are looking for an experienced Site Reliability Engineer working on security infrastructure to design, build, and operate reliable, secure, and scalable infrastructure that underpins identity, access, endpoint, and shared platform services across the company. In this role, you will be a senior technical owner for infrastructure and identity systems end to end, from architecture and implementation through policy enforcement, upgrades, recovery, and day-two operations. You will build durable, production-grade platforms that remove operational friction, enforce security by default, and enable teams to move faster with confidence. This role is well suited for a hands-on senior engineer who thrives in ambiguity, enjoys owning complex systems end to end, and raises the reliability and security bar by replacing fragile implementations with standardized, repeatable infrastructure. This role is based in our San Francisco HQ and requires in-office presence. In this role, you will: Design, build, and operate reliable infrastructure across on-prem, hybrid, shared, and product adjacent environments. Establish standardized infrastructure patterns that replace bespoke implementations with repeatable, auditable, secure-by-default systems. Own the lifecycle of critical infrastructure platforms, including provisioning, deployment, upgrades, patching,
Other cities to consider
More places hiring for this role
Get new service and engagement team leader jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime