About the Team The Finance Platform & Technology team builds and scales the systems and data architecture that power OpenAI’s core financial operations. We enable business agility, compliance, and operational excellence across procure-to-pay, quote-to-cash, supply chain, financial planning, and asset management. We partner with Procurement, Accounting, Tax, Legal, Security, Data, and Engineering to modernize workflows through thoughtful platform design, reliable integrations, scalable automation, and trusted data. About the Role As a Business Systems Lead for Procure-to-Pay, you will be a hands-on engineer who designs, builds, and operates the integrations and first-party applications that power OpenAI’s procurement workflows. You will translate business needs into secure, scalable software, APIs, data flows, and automation across Oracle Fusion, Zip, and connected platforms. You will build the future of buying at OpenAI using OpenAI’s own technology, from guided intake and approval experiences to supplier onboarding, purchasing, receiving, invoicing, and downstream financial data flows. You will own the technical roadmap and support model for these capabilities, improving today’s platforms while deciding where to integrate, configure, or build as OpenAI scales. Your core strength will be software and integration engineering. You will personally write code, troubleshoot cross-system failures, and take solutions through testing, deployment, and production support. You will also make targeted functional configurations in procurement platforms and partner with functional specialists on deeper process and module design. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and operate integrations across Oracle Fusion, Zip, and connected systems using APIs, events, messaging, and batch interfaces where appropriate. Build first-party
Jobs in United States
Production Support Sre Analyst in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current production support sre analyst jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI's mission is to ensure that AGI benefits all of humanity. The Business Systems team helps make that mission possible by building the internal products and platforms that allow OpenAI to operate with speed, reliability, and care. We build internal applications and workflows for Finance and Supply Chain. Our work spans product discovery, React and TypeScript interfaces, Python services and APIs, data models, workflow orchestration, enterprise integrations, and the systems that connect people to systems of record. We work directly with the people who use these products and care about correctness, permissions, auditability, and production reliability. Examples of our work include building an integration platform for supply chain integrations, integrations with Oracle Fusion and Zip, contract intelligence applied to B2B revenue recognition, and Temporal-based agentic workflows for credit checks, duplicate bank detection, and invoice triaging. We turn these efforts into reusable patterns that can support many workflows, rather than one-off automations. About the Role We are looking for Product Engineers to build internal applications end to end. This role spans product discovery, user experience, frontend, backend services, data models, workflow orchestration, and integrations with order management, fulfillment, and supply chain systems. You will take a problem from a first conversation with a Finance or Supply Chain partner through design, implementation, rollout, and production support. Strong candidates combine product judgment with engineering depth. You should be comfortable moving between a React interface, a Python API, a durable workflow, and an integration with an enterprise system. You should be able to ship a useful first version quickly while building the foundations for reuse, security, and long-term maintainability. Direct AI experience is helpful, but the core requirement is strong product engineering judgment and reliable execution. I
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set
About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We’re looking for dedicated, experienced, and deeply curious individuals to help solve some of the most complex challenges faced by our customers while building the future of post-AGI support alongside us. In this role, you’ll work directly with customers through support tickets and live calls, troubleshooting high-impact issues and resolving novel, often ambiguous technical problems in one of the fastest-moving environments in technology. As AI adoption rapidly accelerates, the work you do will directly support mission-critical systems being built on OpenAI’s platform, serving as a critical line of defense for customers operating at massive scale. Beyond resolving technical issues, you’ll help define what world-class support looks like in an AGI-driven future. You’ll partner closely with Engineering, Product, and Operations to improve systems, reduce bugs, and elevate the customer experience, while leveraging automation, agents, and our own AI technology to transform how support operates at scale. You’ll be responsible for: Working directly with customers to troubleshoot and resolve their most complex technical issues, including API failures, integration challenges, authentication errors, and production incidents. Providing end-to-end ownership through debugging logs, analyzing system behavior, reproducing issues, and
About the Team OpenAI’s User Operations team shepherds our customers’ adoption of AI and ensures that our customers' product experience is nothing short of exceptional. We are building the very first post-AGI support team. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others, to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from early-stage startups to established global enterprises. About the Role We’re looking for dedicated, experienced, and deeply curious individuals to help solve some of the most complex challenges faced by our customers while building the future of post-AGI support alongside us. In this role, you’ll work directly with customers through support tickets and live calls, troubleshooting high-impact issues and resolving novel, often ambiguous technical problems in one of the fastest-moving environments in technology. As AI adoption rapidly accelerates, the work you do will directly support mission-critical systems being built on OpenAI’s platform, serving as a critical line of defense for customers operating at massive scale. Beyond resolving technical issues, you’ll help define what world-class support looks like in an AGI-driven future. You’ll partner closely with Engineering, Product, and Operations to improve systems, reduce bugs, and elevate the customer experience, while leveraging automation, agents, and our own AI technology to transform how support operates at scale. You’ll be responsible for: Working directly with customers to troubleshoot and resolve their most complex technical issues, including API failures, integration challenges, authentication errors, and production incidents. Providing end-to-end ownership through debugging logs, analyzing system behavior, reproducing issues, and
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is entering its next phase of growth, with a marketing organization that spans developer audiences, product-led growth, and enterprise go-to-market motions. To support this scale, we are building a modern marketing infrastructure that combines: world-class marketing operations a scalable martech ecosystem lifecycle automation and customer data infrastructure high-impact creative production The Head of Marketing Operations & Production will lead the team responsible for powering this system, and reports directly to the Head of GTM Operations. This role will partner closely with Marketing leadership, Sales and Customer Success Operations, GTM Systems, and Postman’s GTM Strategy and Business Intelligence teams to ensure marketing programs translate into measurable revenue impact. What You’ll Do Marketing Technology & Infrastructure Define and own Postman’s marketing technology ecosystem . You will: Develop and implement the roadmap for Postman’s martech stack, including automation platforms, enrichment tools, campaign systems, and lifecycle infrastructure Ensure reliable and scalable integration between
About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ABOUT BASE LABS Base Labs is a research lab pushing the frontier of open-source LLMs. We think intelligence should be democratized, not controlled by a handful of closed labs and we think very few teams are actually positioned to do something about that. Backed by Baseten's training and inference infrastructure, we have the compute, resources, and talent to take on hard problems at the frontier and open-source what we learn along the way. Our mission is to help build a world where intelligence isn't concentrated, but spread across an ecosystem of models that anyone can build on. That mission shapes what we choose to work on, how we work on it, and who we want in the room. We are accepting applications on a rolling basis for our first cohort of Base Labs Fellows, which is expected to start in late September. Apply using this link. BASE LABS FELLOWSHIP OVERVIEW The Base Labs Fellowship is designed to give researchers exposure to what frontier research looks like in industry. We provide funding, mentorship, and full support to our fellows, with the goal of producing rigorous, published research that shapes both the open-source ecosystem and Baseten's technical roadmap. We run multiple cohorts of Fellows each year and review applications on a rolling basis. This application is for cohorts starting in Sept 2026 and beyond. WHAT TO EXPECT 3 months of full-time research from our San Francisco office A dedicated 1:1 mentorship
OpenAI’s charter calls on us to ensure the benefits of AI are distributed broadly and safely. Our Health AI team focuses on expanding access to high-quality medical expertise and aims to set a high standard for deploying AI responsibly in high-stakes domains. Improving health will be one of the defining impacts of AGI. Today, millions of people lack access to reliable medical information, and clinicians around the world face increasing time and resource constraints. We are building AI systems that support patients, clinicians, and health workers, while meeting the highest standards for safety, reliability, and privacy. We are seeking full stack software engineers to help build and scale products used by consumers and care providers globally. You will work closely with product, design, and research teams to ship real systems in a fast-moving, high-impact environment. In this role, you will: Design and build scalable fullstack systems for consumer and enterprise health. Own end-to-end feature development—from early design and implementation through deployment, monitoring, and iteration. Build and maintain data pipelines and services that meet strict privacy, security, and compliance requirements (e.g., HIPAA). Collaborate closely with researchers and safety teams to integrate reliability, evaluation, and guardrails into production systems. Debug, optimize, and harden systems to support high availability, performance, and global scale. Take ownership of ambiguous problems and drive them to practical, high-quality solutions. You might thrive in this role if you: Are deeply motivated by improving health outcomes and expanding access to medical expertise. Are a strong engineer who enjoys building durable, well-designed systems. Have 5+ years of experience writing maintainable, production-quality code. Can operate with high agency—owning problems end-to-end with minimal supervision. Enjoy working in fast-moving, cross-functional teams with engineers, product managers, desi
About the Team OpenAI’s GTM Partnerships team builds a strategic, global partner ecosystem designed to accelerate customer success, enable enterprise AI adoption, and drive durable growth in support of OpenAI’s mission. We work closely across Product, Research, Sales, Legal, Communications, and regional leadership to ensure a cohesive strategy and disciplined execution with our most strategic partners. About the Role We are hiring a Partner Director, Global McKinsey Alliance to lead and scale OpenAI’s relationship with McKinsey & Company. This role will serve as the single accountable executive owner of the alliance globally. You will define the partnership strategy and operating model, translating executive alignment into scaled commercial and transformation outcomes across regions and industries. You will work closely with McKinsey’s global leadership, industry and functional practices, and technology organizations—including QuantumBlack, AI by McKinsey—to develop differentiated enterprise AI offerings, pursue complex transformation opportunities, and help customers move from strategy to production deployment and measurable value. The role carries responsibility for the health, performance, and long-term expansion of the partnership. Success requires senior executive presence, sound judgment, commercial discipline, technical fluency, and the ability to operate effectively across a complex, partner-led organization. In this role, you will: Own the global McKinsey relationship, including alliance strategy, pipeline growth, enterprise impact, deployment scale, and long-term expansion across priority industries. Define and run the partnership’s operating model across executive governance, joint planning, regional execution, opportunity management, resource allocation, and escalation. Serve as OpenAI’s senior executive counterpart to McKinsey’s global leadership and relevant industry, functional, technology, and regional leaders. Develop a differentiated alliance s
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role Modal's LLM inference platform delivers frontier performance for open-source models with best-in-class elasticity and developer experience, made in part possible by our custom runtime with GPU memory snapshots and multi-cloud substrate . We're looking for a leader to own the direction and execution of this platform to continue to establish us as the clear market leader, working closely with customers like Cognition, Doordash, Ramp, and many more. You'll be leading a group of highly talented engineers working on our market-leading LLM inference offering, spanning the serving stack, routing infrastructure, internal agentic optimization platform, and the user-facing product surface area. This is a hands-on leadership role — expect to split your time between technical contribution, product shaping and people management depending on what the team needs. You'll set direct
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders About the Role We're hiring the first Account Managers at Modal. You'll report to the Regional Director of Account Management and be a founding member of the team. This function does not exist yet. There is no playbook, no territory map, no established motion. You'll own a book of business from day one and build the motion at the same time — from fast-moving AI startups to large enterprise teams running critical infrastructure on Modal. This is a commercial role with a revenue target. You'll be measured on retention and expansion across your accounts. While you won't be delivering the technical recommendations and implementation, the work is technical by nature. Our customers are engineers running GPU workloads, inference, and batch jobs in production, and you need to hold your own in those conversations. The profile we're hiring is a technical account manager. You've worked at companies that are deepl
About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models. We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience. You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments. Key Responsibilities Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency. Develop forecasting models for inference demand across products, regions, and model families. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs. Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions. Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps. Communicate technical findings clearly to both engineering teams and executive leadership. Qualifications MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience). 5+ years of experience working in the infrastructure data science space. Strong ex
Other cities to consider
More places hiring for this role
Get new production support sre analyst jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime