ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Marketing Analytics Manager to establish how Marketing at Baseten makes decisions with data. Marketing at Baseten is scaling fast: more spend, more campaigns, more model launches and more inbound. This is a foundational, hands-on role and the first dedicated Marketing Analytics hire. You'll work directly with Demand Gen, Field Marketing, MOps and Product Marketing alongside GTM, Finance, Product and Engineering to stitch together activities and outcomes across the funnel. You'll build the data models that connect acquisition, engagement, activation, product usage and revenue. You’ll design dashboards, tools, semantic layers and plugins that enable teams to answer questions. Along the way, you’ll develop an understanding of how customers use Baseten and with this, define how our systems and product can improve. RESPONSIBILITIES Define how Marketing success is measured: establish metrics across audience growth, acquisition, activation, engagement, usage, pipeline and revenue. Build the marketing data foundation: ingest and model data across Salesforce, HubSpot, Google Analytics, advertising platforms, web, email, product and third-party sources. Create a medallion architecture that connects the prospect and customer journeys across systems, from first touch through signup, onboarding, activation and expansion. Understand our audiences: develop audience and segmentation frameworks based on customer a
Jobs in United States
Infrastructure And Mlops Engineer in San Francisco
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current infrastructure and mlops engineer jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We’re seeking a senior lawyer to participate in commercial legal strategy and execution across OpenAI’s fast-growing infrastructure portfolio. This is a cross-functional role that will partner closely with procurement, supply chain, partnerships, finance, and product teams to structure, negotiate, and manage the transactions that will support OpenAI’s long-term infrastructure ambitions. We’re looking for an experienced infrastructure transactions lawyer who thrives in ambiguity and wants to help define the commercial playbook for infrastructure efforts in the AI era. This role reports to the Associate General Counsel for infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. In this role, you will: Own commercial legal strategy and risk management for OpenAI infrastructure transactions. Draft, negotiate, and advise on complex agreements with infrastructure suppliers, manufacturers, distributors, and technology partners. Support strategic partnerships involving AI infrastructure and hardware supply chains. Develop frameworks for procurement, licensing, and collaboration across the infrastructure ecosystem. Partner with finance and operations teams to align contract terms with business and compliance requirements. Collaborate with policy and regulatory colleagues on issues impacting global supply chains, export controls, and manufacturing. Build scalable, efficient contracting process
About the Team Our infrastructure team helps deliver OpenAI’s most capable models and products to the world by scaling infrastructure and turning demand into useful FLOPS. We collaborate across research, engineering, design, and business to turn cutting-edge AI advancements into impactful, real-world applications. Our team ensures the right compute is available—at the right time and place—to support some of the world’s most demanding workloads. We empower all of OpenAI’s products and research by scaling the infrastructure behind them. Our work makes it possible to launch new models and products reliably and at scale. About the Role As a Data Scientist on the Infra team, you will play a key role in shaping how we scale the infrastructure that powers OpenAI’s products and research. This is critical as we operate one of the largest and most advanced compute fleets in the world, supporting millions of users and businesses globally. We focus on aligning infrastructure measurement, planning, scaling, allocation, and efficiency to drive measurable impact across the company. You should expect to guide the definition of foundational datasets for infrastructure resources, develop metrics that inform key decisions, build forecasting and optimization models, and establish source of truth dashboards and analyses that enable teams to understand and improve infra usage. Most importantly, you should expect to be a core partner to engineering, research, and product teams in shaping the infrastructure that powers everything OpenAI builds. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain foundational datasets and metrics that reflect infrastructure usage, efficiency, and scaling. Develop forecasting and optimization models to support infra planning and resource allocation. Partner with engineering, research, and product teams to shape infrast
Technical Program Manager – Applied Infrastructure About the Team The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety. The Applied Infrastructure TPM team partners across engineering to lead foundational programs that ensure OpenAI’s infrastructure can meet current and future demand. About the Role We’re looking for a seasoned Technical Program Manager to drive critical infrastructure programs across the Applied organization. This TPM will focus on cross-cutting initiatives such as general compute capacity planning, process transformation, cost and quota attribution and optimization, and coordination across infrastructure and product stakeholders. There will also be focus on evolving OpenAI’s infrastructure to support growth, scale and new products. This work is core to how OpenAI manages and grows its infrastructure footprint in a disciplined, scalable way. Location: San Francisco, CA (Hybrid – 3 days/week in-office) In this role, you will: Serve as the DRI for complex infrastructure programs spanning CPU planning, orchestration, and other resource management domains (e.g. networking, storage). Build and operationalize systems to capture demand signals, model future capacity needs, and align infrastructure planning across internal teams and partners external to the company. Partner closely with Infrastructure, Product and Finance teams to forecast infrastructure usage patterns and ensure supply/demand alignment. Lead cost attribution and quota enforcement programs to promote stability and ensure equitable access to resources across teams. Drive simplification and standardization of infrastructure tooling and processes across Applied and Infra organizations. Drive cross functional programs to evolve our infrastructure to support new growth and scale Work with external v
About the Team The Agent Infrastructure team at OpenAI is responsible for building systems that enable training and deployment of highly useful AI agents, both internally and for the world. We work hand-in-hand with researchers to design and scale the environment in which agentic models are trained – providing a workspace for AI models to execute code, debug issues, and develop software just as human SWEs do. Our training environment for agentic models operates at an extremely high scale and has the flexibility to emulate any environment in which an agent might work. At the same time, our team builds and maintains OpenAI’s core platform for the deployment and execution of agents in production. Our systems power products such as Codex, Operator, tool use in ChatGPT, and future agentic products. Some of the most challenging technical problems in scaling the capabilities and utility of agents and agentic models lie in the infrastructure layer – and our team is focused on building the research and production systems that enable OpenAI to train the most capable models in the world, and maximize the utility of our agentic products for users around the world. About the Role As a Software Engineer on the Agent Infrastructure team, you will have the opportunity to work closely with both research and product at OpenAI - building and scaling systems to train highly capable agentic models, and building the platform and integrations to launch new agents to hundreds of millions of users worldwide. Your work will consist of both building new capabilities - standing up the infrastructure and integrations needed to train more complex agentic models - and rapidly scaling these new capabilities to some of the largest compute clusters in the world. At the same time, you’ll be instrumental to the launch of agentic products at OpenAI - building, maintaining, and scaling the production platform on which all agents run. We’re looking for people with deep experience building AI infrastructu
About the Team OpenAI’s Strategic Finance organization provides the financial insight and guidance that support the company’s long-term strategy and ambitious growth. We partner across Product, Partnerships, Engineering, and GTM to allocate resources to the highest-impact opportunities while preserving sustainable unit economics. The Applications Technology Strategic Finance team works closely with Product and Engineering to scale core offerings and bring new products to market. We lead revenue forecasting and performance management across the product portfolio, and partner closely with the Infrastructure Engineering team to forecast infrastructure and compute demand and to manage margin targets and headcount planning with rigor. About the Role We are hiring a Director of Product Finance, Infrastructure and Compute Demand to help drive strategic decision-making across our product organization. This person will serve as the primary finance partner to the AGI Deployment (Applications) Infrastructure Engineering organization and play a critical role in connecting product usage, compute demand, infrastructure costs, and margin outcomes across the portfolio. You will help build the frameworks that translate technical decisions into economic outcomes, shape planning across product and infrastructure teams, and ensure OpenAI is scaling its products with rigor and efficiency. You will also manage a small team, providing mentorship and leadership while remaining hands-on with analysis and execution. This role is ideally based in our San Francisco HQ, but we are open to NYC and Seattle. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Serve as the primary finance partner to the AGI Deployment (Applications) Infrastructure Engineering organization. Own the consolidated compute and infrastructure demand forecast for the products we bring to market, partnering closely with capacity planning and
About the Team: The OpenAI API team builds the foundation that enables every developer to harness OpenAI’s models safely, reliably, and at scale. We design and operate the systems that power model serving, API access, billing, developer tooling, and enterprise integrations—forming the connective tissue between OpenAI’s research breakthroughs and real-world products. Our mission is to make it effortless for anyone to build with OpenAI technology. We’re responsible for the infrastructure and product layers that allow millions of developers to integrate GPT models, fine-tune behavior, manage data, and deliver transformative experiences to their users. We collaborate across product, research, and engineering teams to ensure that innovation in model capabilities translates directly into value for customers. The API team spans multiple disciplines, including product management, infrastructure engineering, developer experience, and data systems. We care deeply about reliability, scalability, and simplicity—creating tools that let developers focus on their ideas while we handle the complexity of running world-class AI systems. About the Role: We are seeking an experienced Product Manager to define and scale the construction of our data processing, data privacy, billing, and access controls products. You will set strategy and execute on projects like expanding our regional data processing footprint, enabling new inference caching controls in the API or building APIs that make it easier for organizations to manage their spend limits. You will also define the strategy and ship foundational capabilities that ensure customers use OpenAI products securely, privately, and with enterprise-grade controls. This role partners deeply with engineering, security, legal, compliance, finance and leadership to deliver high-trust, enterprise-grade systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to n
About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect our model weights, customer data, and critical systems across multiple cloud environments. We partner with teams across OpenAI—Applied Engineering, Research, IT, and Security—to provide a secure and scalable platform for permissioning, orchestration, and innovative AI research. About the Role We’re looking for a Staff+ Software Engineer to help build and evolve the identity infrastructure that supports OpenAI’s research, engineering, and internal platforms. This role sits at the intersection of cloud infrastructure, identity systems, and software engineering. You’ll work across production systems, infrastructure-as-code, cloud control planes, identity providers, and operational infrastructure to build secure, scalable, and reliable systems used broadly across the company. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time. In this role, you will: Lead the architecture, development, and operation of identity infrastructure that spans cloud platforms, internal systems, and critical engineering services. Design and evolve systems for authentication, authorization, access governance, auditability, and policy enforcement with a strong focus on reliability, scalability, and secure-by-default design. Build foundational infrastructure and platform capabilities that are broadly used across engineering, research, and security teams. Improve the reliability, observability, performance, and op
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As a Member of Technical Staff on AI Infrastructure, you will build and maintain the foundational systems and distributed infrastructure that power AI model post training, inference, and data pipelines. You will collaborate with engineering and research teams to ensure performance, scalability, and reliability of critical AI systems. What You’ll Do Design and implement large-scale, distributed AI infrastructure and services Optimize performance for GPU/xPU accelerators and cloud environments Build tools for observability, reliability, and scaling of AI workloads Partner with cross-functional teams to define AI infrastructure requirements and roadmap Contribute to architectural design and system longevity About You Have experience with GenAI infrastructure systems, distributed systems, cloud computing, and high-performance infrastructure Are proficient in programming languages like Python, Go, or similar Understand scaling challenges specific to AI workloads and accelerators Thrive in fast-paced, collaborative engineering environments The reasonably estimated base salary for this role ranges from $256,000.00 to $276,
£215K – £260K/yr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role Cohere is seeking a Global Public Policy Manager to lead policy engagement on compute infrastructure, export controls, AI competitiveness, and sovereign AI strategies. Governments increasingly view AI as critical national infrastructure and are investing heavily in compute capacity, energy resources, and domestic AI ecosystems. This role will help position Cohere as a trusted partner in emerging discussions around AI infrastructure, national competitiveness, and sovereign AI deployment. Key Responsibilities Monitor and analyze developments related to AI infrastructure, data centers, energy policy, semiconductor policy, export controls, and national AI strategies. Develop policy positions on sovereign AI, compute access, digital sovereignty, and AI competitiveness. Support engagement with governments developing AI infrastructure investment programs and national AI initiatives. Collaborate with commercial, product, and corporate development teams on strategic opportunities involving public-private partnerships. Represent Cohere in policy discussions related to AI infrastructure, energy requirements, and technology c
About the Team The Infrastructure Engineering function sits within IT and is responsible for reliably building, deploying, and operating critical on prem and hybrid environments that power internal services and critical R&D environments. This is an early, high-leverage technical role focused on applying strong Site Reliability Engineering discipline to environments where uptime, safety, recoverability, and security are non-negotiable. This person helps replace bespoke, one-off infrastructure with standardized infrastructure-as-code building blocks that compound reliability and operational leverage as OpenAI scales. About the Role We are looking for an experienced Site Reliability Engineer working on security infrastructure to design, build, and operate reliable, secure, and scalable infrastructure that underpins identity, access, endpoint, and shared platform services across the company. In this role, you will be a senior technical owner for infrastructure and identity systems end to end, from architecture and implementation through policy enforcement, upgrades, recovery, and day-two operations. You will build durable, production-grade platforms that remove operational friction, enforce security by default, and enable teams to move faster with confidence. This role is well suited for a hands-on senior engineer who thrives in ambiguity, enjoys owning complex systems end to end, and raises the reliability and security bar by replacing fragile implementations with standardized, repeatable infrastructure. This role is based in our San Francisco HQ and requires in-office presence. In this role, you will: Design, build, and operate reliable infrastructure across on-prem, hybrid, shared, and product adjacent environments. Establish standardized infrastructure patterns that replace bespoke implementations with repeatable, auditable, secure-by-default systems. Own the lifecycle of critical infrastructure platforms, including provisioning, deployment, upgrades, patching,
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and AI engineers and you'll set the standard for what product looks like here. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and just shipping great product experiences. The role Once a model is deployed, keeping it fast, reliable, and economical at scale is where production inference is won or lost. You'll own the surface that makes that happen: how deployments autoscale, how traffic is routed, how the system fails over, and how workloads scale across clusters and regions. You'll own these as products end to end - both how they work under the hood and how customers configure and observe them - and you'll help set and define the roadmap that infrastructure and product teams alike can build towards. This space is largely still evolving - think Cloud Infrastructure in mid-2000s. Your job is to make it 10x easier to reliably scale and serve AI models in production and set the market standard. Impact and outcomes you'll drive You will own how workloads scale and where they land — autosca
Other cities to consider
More places hiring for this role
Get new infrastructure and mlops engineer jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime