ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a customer-obsessed software engineer to come ship with us. You’ll own features like multi-node training and products like serverless reinforcement learning (RL) from conception to MVP (and from MVP to GA!). You’ll work through the stack, architecting solutions from API and UI down to our infrastructure layer. You’ll fine tune models yourself to develop an understanding of user workflows. You’ll work closely with research engineers leveraging state-of-the-art training techniques to build experiences that accelerate model development and solve for real pain points. If you’re excited to dive deep into the training, let’s talk! THE PRODUCT Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done EXAMPLE INITIATIVES Checkpointing Pipeline: Our checkpointing pipeline starts with automated checkpointing, a feature that ensures that versions of models created during training are automatically backed up to the cloud. Users are able to then deploy checkpoints seamlessly into inference servers, providing point-and-click integrations into inference frameworks like vLLM and Baseten’s Inference Stack. This enables customers to quickly evaluate the performance of their checkpoints with real traffic. Multinode training: Multinode training enables customers to easily run training jobs across multiple compute nodes, enablin
Jobs in United States
Partner Strategy And Operations Lead in San Francisco
1,034 active opportunities · Updated October 2026
Showing
15 jobs
Explore current partner strategy and operations lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and AI engineers and you'll set the standard for what product looks like here. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and just shipping great product experiences. The role Once a model is deployed, keeping it fast, reliable, and economical at scale is where production inference is won or lost. You'll own the surface that makes that happen: how deployments autoscale, how traffic is routed, how the system fails over, and how workloads scale across clusters and regions. You'll own these as products end to end - both how they work under the hood and how customers configure and observe them - and you'll help set and define the roadmap that infrastructure and product teams alike can build towards. This space is largely still evolving - think Cloud Infrastructure in mid-2000s. Your job is to make it 10x easier to reliably scale and serve AI models in production and set the market standard. Impact and outcomes you'll drive You will own how workloads scale and where they land — autosca
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Baseten’s Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost‑efficient. As part of this team, you’ll focus on Model APIs — the infrastructure powering our hosted API endpoints for the latest open‑source models. This work spans distributed systems, model serving, and developer experience. You’ll join a small, high‑impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale. RESPONSIBILITIES: Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations Productionize performa
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. At Baseten, we are building the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA a first-class building block in our infrastructure and unlocking the next generation of distributed inference optimizations. THE OPPORTUNITY Networking and compute are no longer separate disciplines; they are converging. The massive throughput of H100, B200, and NVL72 architectures enables and demands a new approach where communication is co-optimized alongside computation. We are entering an era where the network is an active accelerator, leveraging smart hardware offloads and direct interconnects to ensure that data movement operates at wire-speed. In this role, you will go beyond network configuration to architect the software fabric that unifies thousands of GPUs into a cohesive operating system. While you will leverage the best of the open-source ecosystem, you won't be limited by it. Where off-the-shelf solutions stop, you will build from scratch, engineering the primitives required to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and lightening cold starts. WHAT YOU'LL DO Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE This role sits at the frontier of our research agenda. You will pursue open problems at the intersection of post-training methodology and performant inference, and then collaborate with research engineering to translate findings into production systems. A meaningful portion of your time will be dedicated to research that deepens our understanding of how models learn, alignment, and architectural efficiency — questions that may not have immediate product application. The remainder will be directed toward research that solves concrete problems for Baseten's platform and customers, who are the fastest growing AI companies in the world like Cursor, Lovable, and Notion. We are looking for someone with sharp research taste and genuine creative instinct for problem selection. Someone who can identify questions that matter, design clean experiments to answer them, and push the state of the art. The environment here is not theoretical, but rather research that can be validated with eager customers who are serving billions of tokens a second. RECENT RESEARCH Towards infinite context windows: neural KV cache compaction Dense, on-policy or both? Repeated kv cache for long-running agents Distillation without the dark – replicating black-box on-policy distillation on Baseten RESPONSIBILITIES Define and pursue a research agenda spanning both foundational and applied work, with the applied component connected to Baseten's pla
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. Head of Services & Success Operations Customer Solutions · Reports to Head of GTM Operations · San Francisco, CA (Hybrid) · Full-Time About Postman Postman is the world's leading API platform, with 45M+ developers and 500K+ organizations — including 98% of the Fortune 500. We run a dual motion: bottoms-up product-led growth and a top-down enterprise sales org. The tension between those two motions — and the opportunity to design the operating model that bridges them — is exactly what makes this role interesting. The Customer Solutions Org Customer Solutions runs four functions: Customer Success — developer-forward CSEs; technical depth over relationship management Professional Services — accelerating outcomes through repeatable, monetized programs Solutions Engineering — technical depth in the sales cycle; handoff design is part of this ops role Product Support — inbound; health signals that feed risk models This role is the operational connective tissue across all four. The CS Transformation We're Running We're moving from a traditional enterprise CS model built on relationship management and QBRs to a developer-forward motion
About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for an exceptionally experienced, hands-on full-stack engineer to define and build the next generation of AI-powered enterprise workflows. You will take on the hardest and most ambiguous problems in the Technology vertical: translating real customer needs into product direction, designing the systems behind the experience, and personally writing and shipping production-quality code across the stack. You will own the technical direction and end-to-end delivery of products spanning ChatGPT Work surfaces, backend services, plugins, connectors, enterprise data, permissions, and evaluations. You will make foundational architecture and product tradeoffs; establish patterns other engineers can build on; and hold these experiences to a high bar for reliability, security, observability, and customer value. This is an individual-contributor role for an engineer who leads through technical judgment, direct execution, and influence—not people management. You should be equally comfortable working directly with customers, setting direction with senior cro
About the Team OpenAI for Financial Services is part of OpenAI's Verticals organization, which focuses on accelerating the economy and knowledge work. We build AI products for financial institutions and the professionals who power them, from investment bankers and research analysts to investors and other financial services teams. We combine OpenAI's models, financial data, and enterprise knowledge to help professionals research companies, analyze markets, and produce high-quality work in the tools they use every day. We're a small, entrepreneurial team working closely with customers and partners across research, product, design, engineering, and go-to-market to bring new capabilities from idea to production. About the Role We're looking for full-stack engineers to build new, AI-native products on top of ChatGPT Work and Codex. This is zero-to-one work: you'll help define how financial professionals research, analyze, and make decisions alongside AI. You'll own the experience across the stack, work directly with customers to understand their workflows, and collaborate across OpenAI to turn new model capabilities into products that professionals can trust. Your work will shape how some of the world's largest financial institutions adopt AI and how financial knowledge work gets done. In this role, you will: Build a new financial services app within ChatGPT Work and Codex, creating intuitive AI-native experiences for company research, financial analysis, document review, and professional work products. Develop new product experiences around enterprise memory that learn from an organization's knowledge, workflows, and context, and adapt to how its teams work. Develop the APIs, services, and integrations required to connect user experiences with financial data providers, enterprise systems, and OpenAI's models. Work closely with product and design to turn ambiguous customer problems into polished, useful, and reliable products. Work directly with financial institutions to
About the Team OpenAI for Financial Services is part of OpenAI's Verticals organization, which focuses on accelerating the economy and knowledge work. We build AI products for financial institutions and the professionals who power them, from investment bankers and research analysts to investors and other financial services teams. We combine OpenAI's models, financial data, and enterprise knowledge to help professionals research companies, analyze markets, and produce high-quality work in the tools they use every day. We're a small, entrepreneurial team working closely with customers and partners across research, product, design, engineering, and go-to-market to bring new capabilities from idea to production. About the Role We're looking for backend engineers to build the systems that make advanced AI useful, reliable, and trustworthy in financial services. You'll build the data systems, agentic workflows, and enterprise integrations behind our products. You'll also help bring them into production at some of the world's largest financial institutions. This is a product-minded engineering role with significant ownership and zero-to-one building. You'll shape new products from the ground up, work directly with customers to understand their workflows, and collaborate across OpenAI to turn new model capabilities into products that professionals can trust with high-stakes work. In this role, you will: Design and build backend systems that power AI-native financial workflows across ChatGPT Work and Codex. Build infrastructure to ingest, index, retrieve, and serve financial data, company filings, market information, and firm-specific knowledge at scale. Develop integrations with financial data providers, enterprise knowledge systems, and customer environments, including the authentication, authorization, and entitlements required to use them securely. Build the systems that let models and agents use the right tools and data, preserve source provenance, and produce accurate,
About the Team OpenAI’s Global Affairs team engages with governments, policymakers, civil society, community organizations, and other stakeholders around the world. The State and Local Government Affairs team leads OpenAI’s engagement with state and local governments across the United States and works cross-functionally to advance policies and partnerships that support OpenAI’s mission. About the Role As Government and Community Affairs Manager, you will help support OpenAI’s engagement with state and local policymakers, community organizations, civic leaders, educational institutions, workforce partners, and other stakeholders across California. Reporting to the State and Local Government Affairs Lead, Western Region, you will develop and maintain relationships with elected officials, agency leaders, legislative staff, local government representatives, and community partners. You will monitor and analyze policy developments, support advocacy and engagement strategies, and help ensure that California stakeholders understand OpenAI’s technology, mission, and approach to responsible AI development. This role extends beyond traditional government affairs. You will also help develop community partnerships, educational initiatives, events, and other programs that demonstrate how AI can benefit people and communities across California. You will work closely with colleagues across Global Affairs, Legal, Communications, Partnerships, Product, Go to Market, and other teams to connect OpenAI’s policy priorities with meaningful external engagement. Although the role will focus primarily on California, you may also support engagement in other Western states on an ad hoc basis. We are looking for someone who is flexible, collaborative, comfortable operating in a fast-moving environment, and able to approach emerging needs with sound judgment and a bias toward action. In this role, you will: Help develop and execute OpenAI’s state and local government affairs and community engage
About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for a full-stack product engineer who can own ambiguous enterprise workflows end to end: understand a customer problem, shape the product, build across frontend and backend, work through platform dependencies, instrument quality, and learn quickly with design partners. You will build across ChatGPT Work surfaces, services, plugins, connectors, and data or permission boundaries when the experience requires it. You will make quality and rollout observable through evaluations, product and operational signals, and clear fallback or rollback paths. This is a product-engineering role for someone who can move between user problems and system details without losing ownership of either. The strongest candidates will be able to turn specific customer evidence into a generalizable product, explain scope and architecture tradeoffs, and carry a feature from an early prototype through a bounded production rollout. In this role, you will: Build and ship role-specific workflows across ChatGPT Work surfaces, services, plugins, and connectors. Turn customer a
About the Team At OpenAI, our Trust, Safety & Risk Operations teams safeguard our products, users, and the company from abuse, fraud, scams, regulatory non-compliance, and other emerging risks. We operate at the intersection of operations, compliance, user trust, and safety working closely with Legal, Policy, Engineering, Product, Go-To-Market, and external partners to ensure our platforms are safe, compliant, and trusted by a diverse, global user base. The Global Safety Response Operations team within the org provides 24/7 coverage for user safety, risk, and regulatory escalations across OpenAI’s products, handling the highest-priority cases that require human judgment and rapid response. The team operates as the core escalations management and delivery arm of OpenAI’s safety operations, ensuring that our products remain safe and aligned with our policies while enabling timely, empathetic, and consistent user support. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Please note: This role may involve exposure to sensitive content, including material that is sexual, violent, or otherwise disturbing. About the Role We’re looking for experienced Trust, Safety, and Risk Operations analysts who have subject matter expertise in one or more of the following areas: policy enforcement and content moderation, fraud and scam prevention, developer risk, or privacy and regulatory escalations. You’ll be on the front lines of safety escalation management, helping to triage and resolve urgent and sensitive cases. You’ll work across subject matter areas, systems, and processes to ensure operational excellence, develop process improvements and automations, and surface insights and trends. This is a 24/7 global operation that requires flexibility to work rotating shifts, including nights, weekends, and holidays, as part of an on-call coverage model. We use a hybrid work model of 3 days in the office per week and offer r
About the Team Critical Harm Operations sits within User Safety & Risk Operations and builds enforcement systems for Frontier Risk and Material Harm that are accurate, fast, defensible, and built to scale. We turn policy intent into operational readiness, review standards, quality systems, escalation paths, automation guardrails, and durable cross-functional operating models. About the Role We are looking for an exceptional Program Manager to help build durable operating systems and run some of OpenAI’s most complex safety operations. The core need is a high-agency operator who can take an ambiguous problem, create the right structure, align cross-functional partners, and drive the work through execution. This role will move across Critical Harm priorities as needs evolve. You may step into operationalizing national security or violent-activities workflows, support wellbeing and Frontier Risk initiatives, or help scale programs such as Trusted Access. Deep domain expertise is helpful but not required; the strongest candidates will learn quickly, exercise excellent judgment, and make complex programs move. In this role, you will: Lead strategic operational builds across priority workflows from problem statement to implemented operating model, including scope, owners, milestones, risks, success measures, and execution cadence. Translate policy, safety, technical, legal, and operational constraints into workflows, requirements, playbooks, escalation paths, and decision-making structures that teams can execute. Coordinate with User Ops leadership, Product Policy, Integrity, Safety Systems, i2, Legal, Product, Engineering, Support, vendors, and other partners to resolve dependencies and keep critical work moving. Move in and out of workflows as priorities shift—standing up new programs, stabilizing operations, improving handoffs, and transitioning durable ownership to the right team. Use operational data and frontline signals to identify bottlenecks, quality gaps, ca
Other cities to consider
More places hiring for this role
Get new partner strategy and operations lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime