About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late
Jobs in United States
Lead Lead Technical Program Manager Infrastructure Consultant in San Francisco
414 active opportunities · Updated September 2026
Showing
15 jobs
Explore current lead lead technical program manager infrastructure consultant jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
Hiring demand
69/100
rising · 105 related jobs
Hiring trend
+76.3%
Job postings compared with the previous 30 days
Remote options
21.9%
Share of matching jobs listed as remote
Typical salary
$190K – $190K/yr
Based on 11 salary observations
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
About the Team OpenAI’s Finance and Revenue Operations organization builds the commercial infrastructure that enables the business to scale with speed, discipline, and financial integrity. Deal Desk operates as the commercial strategy and governance function at the intersection of Sales, Partnerships, Legal, Technical Revenue, Finance, Order Management, Billing Operations, Product, and GTM Systems. We architect complex enterprise transactions, turn ambiguity into executable decisions, and create the guardrails that let the business move quickly with operational and financial discipline. As OpenAI’s enterprise business grows in scale and complexity, Deal Desk defines how novel commercial motions become durable operating capabilities. We convert precedent-setting deal decisions into repeatable policy, controls, workflows, and systems requirements. About the Role We are hiring a Strategic Deals & Commercial Architecture Lead to lead the structuring and governance of OpenAI’s most complex enterprise transactions. This is a senior individual-contributor leadership role for someone with exceptional enterprise deal judgment, operational rigor, and a builder mindset. You will serve as the commercial architect for high-stakes opportunities, translating ambiguous requirements into coherent deal structures, approval strategies, and executable quote-to-cash plans. Your work will shape more than individual transactions. You will establish decision principles and precedent, clarify tradeoffs, and turn recurring patterns into scalable guidance, controls, systems requirements, and enablement. You will own the commercial decision and governance layer that helps strategic opportunities move decisively while managing downstream risk across contracting, billing, revenue recognition, provisioning, reporting, controls, auditability, and customer experience. The right candidate can move seamlessly between advising on a single high-value transaction and improving the operating model be
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role At Sentry, Support is an engineering discipline. Our customers are the greatest technical minds in the world—developers at elite enterprises building the future of software—and they deserve answers that go deeper than a knowledge base link. We're looking for an APAC Technical Support Engineer based in San Francisco to join our global Support Engineering team. This role is designed to provide APAC coverage to our users; with the shift being Sunday through Thursday 4PM-12AM PST. We are architecting the Technical Support engine . We’re looking for an experienced engineer to help us redefine the standard of technical support by combining deep human expertise with autonomous agentic systems. You are a debugger of both code and systems. You will treat support volume as a data signal to build automated resolution paths, ensuring our human engineers only touch the most complex, high-impact architectural puzzles. Sentry Support Engineers aren't just clearing queues; they are Orchestrators . You will engage with our users across GitHub, Discord, and our internal systems, while acting as the Technical Lead for our Agentic Ops. You ensure that when a developer asks a complex question, our systems have the right context and a seamless "Human-in-the-Loop" path to you when deep, nuanced expertise is required. In this role you will Master the Sentry Ecosystem & Support Elite Developers Deep-Dive Debugging: Perform root-cause analysis on complex issues and distributed tracing gaps across polyglot environments. Support the Great Minds: Act as a strategic consultant for senior engineers at our largest enterprise customers, s
What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.
About the team The AI Deployment Engineering team is responsible for helping developers and enterprises safely and effectively deploy OpenAI technologies in production. We act as trusted technical advisors and thought partners for customers, working side by side with their teams to identify high-value use cases, design practical architectures, and move from prototype to durable deployment. Cybersecurity is one of the most urgent domains where AI can help. Security teams are under pressure to reason across code, logs, infrastructure, tickets, alerts, and vulnerability data faster than ever. As frontier models become more capable, organizations need deep technical guidance on how to evaluate, validate, and safely deploy AI systems in security-critical workflows. About the role We are looking for a Cyber AI Deployment Engineer to partner with customers and help them apply OpenAI models, APIs, Codex, and agentic workflows to real cybersecurity use cases. You will work with CISOs, security executives, application security leaders, SOC teams, security engineering teams, and hands-on practitioners to identify where AI can create measurable security outcomes. This is a customer-facing technical role for someone who can move fluidly between executive strategy, practitioner-level cyber depth, and hands-on solution design. You will help customers evaluate and deploy workflows such as secure code review, vulnerability triage, threat modeling, remediation, SOC and incident response workflows, detection engineering, cloud security, GRC automation, and security validation. You will collaborate closely with Sales, Solutions Engineering, Product, Engineering, Research, and Security to turn customer needs into safe deployment patterns, reusable field assets, and product feedback. This role is based in our San Francisco HQ. We offer relocation support to new employees. In this role, you will: Deeply embed with strategic customers as the technical lead for AI-enabled cybersecurity work
From $177.2K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We're looking for a Staff Software Engineer to lead the technical direction of the backend systems powering Pinterest's AI-driven product experiences — Pinterest Assistant, visual editing and content creation tools and future LLM-based products. You'll design and ship backend systems while architecting the broader platform strategy that enables these experiences to scale across surfaces and teams. This is a hands-on leadership role where you'll move between deep technical execution, system-level architecture and cross-team technical leadership. What you'll do: Define the backend and platform architecture for AI-driven product experiences — visual-chat, AI image generation and editing, and agentic or LLM-based products — partnering with Engineering, Product, ML and UX leaders to shape the technical vision and roadmap. Architect end-to-end systems
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are the Data Foundation & AI team within Plaid’s Data organization. Our mission is to build the shared ML and AI infrastructure that powers intelligent capabilities across Plaid’s product suite. We develop the foundational systems, models, and data assets that transform Plaid’s unique financial network data into scalable, general-purpose representations that teams across the company can leverage. Our work spans the full ML lifecycle — from large-scale data curation and model pretraining to production serving, evaluation, and monitoring. As part of the team, you’ll work at the intersection of machine learning infrastructure, applied AI, and distributed systems, helping establish the core AI platform that enables innovation across Plaid. As a Staff Machine Learning Engineer, you will lead the technical strategy and development of Plaid’s foundation models, driving key decisions across pretraining objectives, model architecture, and fine-tuning approaches that power a wide range of downstream product applications. You will serve as the technical lead for the full machine learning lifecycle, overseeing everything from data curation and experimentation to production deployment, feature management,
From $268.1K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We are looking for a Sr. Staff Machine Learning Engineer to be the Technical Lead for the Content Quality who will build the overall technical strategy, unified technical architecture and define a roadmap for industry leading methodology. We are seeking strong hands on machine learning background including content modeling, signal lifecycle, and platforms used to enforce signal use with downstream use cases. You’ll be working with other leads to set and execute a long-term strategy for the team, aligning the strategy with other clients where it makes sense and communicating to leadership our current status and path to having world-class capabilities. You'll also foster a healthy community where all Content Quality engineers can learn best practices, collaborate effectively and understand our technical direction. What you’ll do: Arch
About the team The Agent Enablement AI Deployment Engineering (ADE) team works across engineering, product, design, partnerships, and strategic customers to grow an open ecosystem of agent-enabled sites and services. We help partners adopt the OpenAI tech stack related to identity, permissioning, agent-auth primitives so users can safely connect ChatGPT and Codex to the tools, services, and workflows they already use. Our team also works with external partners on defining the standards for agent access, marketplace offerings as well as other agent enablement initiatives to ensure users of ChatGPT and Codex go from intent to task completion seamlessly. About the role We are looking for an AI Deployment Engineer to help strategic partners design, build, validate, launch, and operate agent enablement integrations across web applications, connectors, APIs, CLIs, MCP servers, and developer tools. This is a hands-on, partner-facing product engineering role for someone who can contribute to the platform itself, lead sophisticated technical engagements, and turn ambiguous identity and agent-workflow requirements into secure, production-ready integrations. You will work across partner product and engineering teams and OpenAI’s product, engineering, design, partnerships, legal, policy, security, support, and go-to-market teams. You will identify high-value user journeys, choose the right integration path, prototype and review architectures, write code, run evaluations and dogfood, trace failures end to end, guide launch and rollout, and support post-launch iteration. The best person for this role moves fluidly between full-stack code, OAuth/OIDC and identity systems, product judgment, project leadership, and clear communication with engineers and executives. This role is a fit for a product-minded engineer who wants to stay close to users and partners while going deep on authentication, permissions, reliability, safety, and developer experience. The principle objective is to
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by leading the design of high-performance, scalable and reliable machine learning systems? Do you want to set technical direction and help shape the next generation of AI platforms powering advanced NLP applications? We are looking for a Lead Member of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team. You may be a good fit if you have: 8+ years of engineering experience running production infrastructure at a large scale, with a track record of technical leadership Demonstrated experience leading the architecture
About the Team OpenAI’s Forward Deployed Engineering (FDE) team turns research breakthroughs into production-grade systems. We embed deeply with customers to solve high-leverage problems and act as the delivery engine for our most complex large-scale engagements. We move quickly from prototype to production and surface reusable patterns that shape our platform. We operate at the intersection of deployment and development – working closely with OpenAI Research, Product and Partnerships. About the Role As a Technical Deployment Lead (TDL), you will define how OpenAI delivers complex systems to Semiconductor customers. You will own how solutions are scoped, built, shipped, and adopted across high-value engineering workflows such as RTL design, verification, and physical implementation. You’ll translate business outcomes into a technical plan, run day-to-day execution across FDEs, Researchers, and Customer Engineers, and partner with customer teams to ensure delivery supports their goals. You will focus on the semiconductor vertical to deploy next-generation AI capabilities. You will own delivery end-to-end: embedding with customers to map workflows and success criteria, ensuring components ship on time, and leading readiness and change management for adoption. You’ll track progress, manage dependencies, make sequencing decisions, and drive 0→1 prototypes through MVP and scale. You will also share field insights with Product and Research to guide roadmap and priorities. Success will be measured first and foremost by impact - deployments that deliver measurable value against customer goals, drive adoption, and become critical to their workflows. Additional measures of success include delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact (field signal that shifts roadmaps/architectures). This is a high-trust, high-autonomy role. Success requires deep technical project m
About the team OpenAI’s Forward Deployed Engineering (FDE) team turns research breakthroughs into production-grade systems. We embed deeply with customers to solve high-leverage problems and act as the delivery engine for our most complex large-scale engagements. We move quickly from prototype to production and surface reusable patterns that shape our platform. We operate at the intersection of deployment and development – working closely with OpenAI Research, Product and Partnerships. About the Role As a Technical Deployment Lead (TDL), you will define how OpenAI delivers complex systems to customers. You will own how they are built, shipped, and adopted. You’ll translate business outcomes into a technical plan, run day-to-day execution across FDEs, Researchers, and Customer Engineers, and partner with customer teams to ensure delivery supports their goals. You will own delivery end-to-end: embedding with customers to map workflows and success criteria, ensuring components ship on time, and leading readiness and change management for adoption. You’ll track progress, manage dependencies, make sequencing decisions, and drive 0→1 prototypes through MVP and scale. You will also share field insights with Product and Research to guide roadmap and priorities. Success will be measured first and foremost by impact - deployments that deliver measurable value against customer goals, drive adoption, and become critical to their workflows. Additional measures of success include delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact (field signal that shifts roadmaps/architectures). This is a high-trust, high-autonomy role. Success requires deep technical project management expertise, extreme ownership of outcomes, and an ability to immerse in customer workflows and partner with customer teams to solve complex engineering problems at pace. This role is based in San Francisco. W
About the Team OpenAI’s Education team is building products that advance how people learn with AI. The team works across higher education institutions, K-12 districts, and country-level partnerships, including applied research on how AI affects learning and cognitive outcomes. The team owns owns ChatGPT Edu, ChatGPT for Teachers, and related product/research work. The team partners closely with go-to-market, research, Consumer Learning, and model teams to turn education-specific insights into product experiences that can improve ChatGPT more broadly. Some of our recent work: New Education Plugins for ChatGPT Work and Codex New tools for understanding AI and learning outcomes Education for countries Advancements in higher education Early product work - Introducing Study Mode About the Role We’re looking for a hands-on Tech Lead Manager to lead and manage a team of senior full-stack engineers building AI-native learning experiences in ChatGPT. This person will combine technical execution, product judgment, and people leadership: they will write and ship code, manage engineers, and help shape the product direction for how students and Educators use AI. In This Role, You Will Lead and manage a team of three senior full-stack engineers. Build product experiences for ChatGPT Education, ChatGPT for Teachers, and AI-native learning workflows. Partner with research teams on field studies, randomized control trials, classifiers, data pipelines, and cognitive-outcome measurement. Collaborate with Consumer Learning and model teams to translate education insights into broader ChatGPT behavior and product improvements. Drive execution across product, engineering, research, go-to-market, and partner teams. Help define product strategy, priorities, and delivery plans for a new product pod. You Might Thrive In This Role If You Have several years of direct people-management experience with engineers. Are still highly technical and comfortable doing IC engineering work. Have strong pr
Higher-paying openings
Jobs with higher listed pay
Technical Enablement Program Manager
Postman · San Francisco, California, United States
$1.6M – $2M/yr
Sales Enablement Program Manager
Postman · San Francisco, California, United States
From $1.4M/yr
Supply Chain Program Manager (SCPM) - AI Infrastructure
OpenAI · San Francisco, California, United States
$226K – $285K/yr
Staff Technical Program Manager
Sentry · San Francisco, California, United States
$200K – $240K/yr
Related career options
Similar roles with stronger pay
Demand 35/100 · 7 jobs
$4.6M – $4.6M/yr
Salary →Demand 46/100 · 8 jobs
$345K – $345K/yr
Salary →Demand 73/100 · 93 jobs
$294.5K – $294.5K/yr
Salary →Demand 51/100 · 22 jobs
$292.5K – $292.5K/yr
Salary →Demand 45/100 · 9 jobs
$255.7K – $255.7K/yr
Salary →Demand 77/100 · 123 jobs
$242.1K – $242.1K/yr
Salary →Other cities to consider
More places hiring for this role
Get new lead lead technical program manager infrastructure consultant jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime