About the team OpenAI’s Forward Deployed Engineering (FDE) team turns research breakthroughs into production-grade systems. We embed deeply with customers to solve high-leverage problems and act as the delivery engine for our most complex large-scale engagements. We move quickly from prototype to production and surface reusable patterns that shape our platform. We operate at the intersection of deployment and development – working closely with OpenAI Research, Product and Partnerships. About the Role As a Technical Deployment Lead (TDL), you will define how OpenAI delivers complex systems to customers. You will own how they are built, shipped, and adopted. You’ll translate business outcomes into a technical plan, run day-to-day execution across FDEs, Researchers, and Customer Engineers, and partner with customer teams to ensure delivery supports their goals. You will own delivery end-to-end: embedding with customers to map workflows and success criteria, ensuring components ship on time, and leading readiness and change management for adoption. You’ll track progress, manage dependencies, make sequencing decisions, and drive 0→1 prototypes through MVP and scale. You will also share field insights with Product and Research to guide roadmap and priorities. Success will be measured first and foremost by impact - deployments that deliver measurable value against customer goals, drive adoption, and become critical to their workflows. Additional measures of success include delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact (field signal that shifts roadmaps/architectures). This is a high-trust, high-autonomy role. Success requires deep technical project management expertise, extreme ownership of outcomes, and an ability to immerse in customer workflows and partner with customer teams to solve complex engineering problems at pace. This role is based in San Francisco. W
Jobs in United States
Platform Engineering Director in San Francisco
778 active opportunities · Updated October 2026
Showing
15 jobs
Explore current platform engineering director jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. Make a Difference: Monitor and maintain deployed m
About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol
Join the engineering teams that bring OpenAI’s ideas safely to the world! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic-testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU, storage, GPU, and network lifecycle management to drive efficiency, accountability and dynamic optimization of our resources. Implement fault-tolerant and resilient design
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About North: North is Cohere's cutting-edge AI workspace platform, designed to revolutionize the way enterprises utilize AI. It offers a secure and customizable environment, allowing companies to deploy AI while maintaining control over sensitive data. North integrates seamlessly with existing workflows, providing a trusted platform that connects AI agents with workplace tools and applications. Why this role? This role offers a unique opportunity to shape how enterprises harness the power of AI in real-world applications. As a bridge between our core North product and our clients’ engineering teams, you’ll be at the forefront of solving complex problems and securely integrating AI into critical sectors such as finance, healthcare, and telecommunications. We’re looking for Software Engineers with Applied AI experience who can own the design, build, and deployment of agentic workflows powered by Large Language Models (LLMs), from early prototypes to production-grade AI agents, to deliver concrete business value in enterprise workflows. You’ll work closely with customers on real-world business problems, often building first-of-thei
About the Team OpenAI’s Platform team powers how millions of developers and enterprises build with our models. We provide APIs and agentic solutions used by global startups and fortune 500s. We work closely with product, engineering, design, and go-to-market to build a world-class platform that pushes the frontier of AI capabilities. About the Role As a Data Scientist on the Platform team, you will drive a data-driven culture for OpenAI’s API and B2B solutions. You’ll define the metrics that matter for developer success and enterprise value, measure the impact of new models and features, and partner with PMs and engineers to improve model quality, reliability, latency, and cost. Your work will shape how thousands of products adopt agentic AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Platform product team as a trusted partner, uncovering ways to improve developer experience, reliability, and usage growth Define north-star metrics across the developer funnel (activation, retention, growth), as well as latency/cost guardrails for new features and models Design and interpret A/B tests and controlled rollouts (e.g., new model versions, pricing/limits, new API features, new B2B products) Build source-of-truth dashboards and self-serve data tools for product, engineering, and go-to-market teams Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality) You might thrive in this role if you have 5+ years in a quantitative role in ambiguous, high-growth environments (platforms, APIs, or B2B products a plus) Depth in SQL and Python, with a track record proposing, designing, and running rigorous experiments Experience defining and operationalizing metrics from scratch (including reliability/latency/cost and safety) Strong cross-functional communication with PMs, enginee
About the Team The Premium team owns some of the highest-leverage customer-facing levers in ChatGPT’s consumer revenue business, spanning the paid customer journey: helping users understand the value of paid plans, convert with confidence, and continue finding lasting value in their subscription. Our work is highly cross-functional, partnering with Product, Data Science, Design, FinEng, Finance, Legal, Support, and Marketing to improve free-to-paid conversion, renewal, customer lifetime value, and revenue while keeping the experience trustworthy, scalable, and low-friction. In This Role, You Will: Lead and scale an engineering team responsible for some of ChatGPT’s most important subscription and monetization experiences. Own the technical execution for Premium customer experiences across plan merchandising, paywalls, upgrade flows, checkout UX, plan management, renewals, downgrades, and cancellation. Partner with Product and Data Science to run high-quality experiments across upgrade, trial, renewal, downgrade, and cancellation flows. Improve key subscription metrics including conversion, renewal, churn, ARPU, and lifetime value. Build reliable customer-facing Premium experiences for purchase, plan management, renewal, downgrade, cancellation, and access-related states at scale. Partner closely with FinEng and other platform teams to evolve the billing, payments, and entitlement capabilities that power Premium experiences. Collaborate closely with Product, Design, Data Science, Finance, Legal, Support, and Marketing on monetization strategy and execution. You Might Thrive in This Role If You: Have 5+ years of engineering management experience, Have strong technical expertise in backend, frontend, or full-stack development, with experience building growth-oriented features. Have a track record of improving conversion, retention, or monetization through experimentation and data-driven product engineering. Are experienced with subscription products, plan merchandising
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The largest, most demanding enterprises are starting to run on Baseten, and they arrive with a range of security, compliance, and procurement requirements. As a Senior Engineer on Baseten's enterprise engineering team, you'll build the capabilities that enable large organizations like Writer, HubSpot, and Notion to succeed on Baseten. Enterprise engineering authors the core building blocks, APIs, and user experiences powering the Baseten platform: identity and access management, billing, regional isolation, and self-hosted and single-tenant deployment options. This is deep product and systems work across the full stack, from designing authentication and authorization systems using standards like OAuth and OIDC to shipping the admin experiences enterprise IT teams use to manage their organization. EXAMPLE INITIATIVES Recent and upcoming work on the team: Fine-grained authorization for users, service accounts, and agentic workloads SSO and SCIM support, allowing customers to centralize and automate access to Baseten Expanding the billing platform to support evolving pricing models, advanced data exports, and controls to manage spend In-product management and enforcement of customer compliance requirements like data residency and HIPAA Securing network paths in and out of a customer's models with private connectivity and ingress and egress restrictions Allowing customers to run Baseten inside their own VPC, on-pr
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Security Engineering is the engineering function inside the Plaid security org that focuses on developing the industry-leading security systems and infrastructure. Security Engineering owns most of Plaid’s security-related infrastructure: secure data storage, key management systems, internal identity platform, internal authentication systems, internal permission management, and internal authorization service. We develop solutions across data encryption, key management, access control, and data loss prevention to protect sensitive consumer data. We believe in the Zero Trust security model and are always looking for ways to improve our authentication and access control platforms. About the role You will develop security capabilities to secure Plaid infrastructure and to secure sensitive data access. You will own, maintain, and build Plaid’s security infrastructure and services like Key Management System and Secure Token Service. You will consult with product engineers to ensure Plaid services meet security standards. You will help educate and support other engineering teams to improve security in their own products and services. You will assist with Plaid’s incident response and security awareness pro
About the Team The Future of Computing Research team is an Applied Research team within the Consumer Devices group focused on developing new methods and models as we advance forward in our mission of building AGI that benefits all of humanity. As a Software Engineer on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. About the Role We are looking for a Software Engineer to join our team to build tools and services that enable AI research, evaluation, and data generation workflows. The best work in this role will start with an ambiguous design question and turn it into working research systems. You will work closely with researchers, designers, and engineers to build the evaluation systems, synthetic data generation pipelines, review tools, and supporting platform services. The goal is to make these workflows easier to create, run, and trust without requiring bespoke engineering support for each new design concept. You will help ensure that research artifacts have a clear lifecycle, runs are reproducible and observable, and results provide useful evidence for product and model-training decisions while the underlying systems remain reliable and reusable. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Build web applications, APIs, data models, and backend services for AI research workflows. Build tools to author and manage evaluation tasks, rubrics, graders, suites, and rollout configurations, including workflows for publishing, versioning, auditing, and sharing research artifacts. Automate evaluation runs and generate useful reports for design, research, and engineering teams. Support synthetic data generation workflows for multimodal and conversational research, including tools that comb
About the Team Our team turns OpenAI’s latest model capabilities into polished, trusted products for consumers and developers. We build the end-to-end experiences including product surfaces, platform layers, and developer workflows that make cutting-edge AI accessible, useful, and dependable at scale. OpenAI’s Financial Engineering (FinEng) team powers how revenue flows through our products - pricing and packaging, checkout, payments, subscriptions, and the financial infrastructure behind them. We partner closely with Engineering, Data Science, Risk, Finance, and Go-to-Market to make paying for OpenAI products seamless, reliable, and efficient worldwide. We pair rapid innovation with a rigorous approach to responsible deployment. Safety and trust are built into how we design, ship, and learn from real-world usage, so these tools deliver meaningful value while aligning with OpenAI’s mission. About the Role We are seeking an experienced Product Manager to scale the product efforts and technical strategy within our Financial Engineering team. The ideal candidate has prior experience in billing, finance, and accounting, ideally also building solutions for commercial users of varying sizes from small scale to enterprise. This role requires close collaboration with our product, finance, operations, and engineering teams. This position is based in San Francisco, CA. We utilize a hybrid work model with 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Develop a strategy and roadmap to efficiently scale the billing operations and customer experience behind OpenAI’s growing product portfolio Identify and execute opportunities to improve the order-to-cash processing for OpenAI’s largest and most strategic customers Build AI powered tooling for key partner teams such as Finance and User Operations to drive better decisions and business outcomes Collaborate with other product teams to defining OpenAI’s evolving monetization s
About the Team The Platform Analytics team builds the systems OpenAI researchers use to understand the quality and behavior of the models we train including what models are doing, why they behave in a particular way, and how that behavior changes across experiments. Neptune is a core part of this work. It ingests, stores, queries, and visualizes large volumes of metrics from pretraining, post-training, and reinforcement learning. Hundreds of researchers depend on these systems in their daily work to compare experiments, debug unexpected behavior, and decide what to try next. Our scope is broader than metrics. We also build platforms that help researchers analyze samples, traces, evaluation results, and other structured or unstructured data through dashboards, APIs, and increasingly agent-driven workflows. These systems need to remain fast, reliable, and understandable as the scale and complexity of research change quickly. We are not trying to become a consulting team that builds a separate solution for every research project. We work directly with researchers to understand recurring problems, then turn them into reusable infrastructure and platform capabilities that many teams can build on. About the Role We’re looking for a hands-on experienced software engineer who can take ownership of a critical system and drive it from problem definition through production adoption. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users. The right candidate should already bring the technical judgment, ownership, and execution expected at this level. The primary learning curve should be OpenAI’s stack and research problem space, not learning how to lead a complex engineering effort or deliver a production system. You will work directly with
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for an experienced Executive Assistant to support our CTO (co-founder) and Head of Engineering. This is a highly operational role that goes well beyond calendar management. You’ll own the day-to-day operating rhythm of the Engineering organization, ensuring leaders are prepared, priorities stay coordinated, and critical meetings, communications, and follow-ups happen seamlessly. You’ll partner closely with senior engineering leaders and serve as a trusted point of coordination for employees, customers, candidates, and external partners. Success in this role comes from exceptional organization, judgment, attention to detail, and the ability to keep many moving pieces aligned in a fast-growing environment. RESPONSIBILITIES Own complex calendar management for the CTO and Head of Engineering, balancing shifting priorities while ensuring time is allocated intentionally Ensure leaders are prepared for every day and every meeting by proactively managing agendas, materials, context, logistics, and follow-ups so time is used effectively and decisions move forward Support forward-looking calendar planning, coordinating recurring operating cadences including roadmap planning, leadership meetings, P0 reviews, Engineering All Hands, and other cross-functional forums Own the operational cadence of the Engineering organization, including weekly leadership meetings, monthly Show & Tells, Engineering All Hand
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. Product at Baseten Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and AI engineers and you'll set the standard for what product looks like here. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and just shipping great product experiences. The role Once a model is deployed, keeping it fast, reliable, and economical at scale is where production inference is won or lost. You'll own the surface that makes that happen: how deployments autoscale, how traffic is routed, how the system fails over, and how workloads scale across clusters and regions. You'll own these as products end to end - both how they work under the hood and how customers configure and observe them - and you'll help set and define the roadmap that infrastructure and product teams alike can build towards. This space is largely still evolving - think Cloud Infrastructure in mid-2000s. Your job is to make it 10x easier to reliably scale and serve AI models in production and set the market standard. Impact and outcomes you'll drive You will own how workloads scale and where they land — autosca
Other cities to consider
More places hiring for this role
Get new platform engineering director jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime