Jobs in United States

Ai Platform Director in San Francisco

1,474 active opportunities · Updated October 2026

Explore current ai platform director jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Strategic Finance team partners across OpenAI to turn company strategy into financial decisions, helping allocate resources and steer the business toward its highest-impact long-term outcomes. Within Strategic Finance, the B2B Product Finance team owns the financial perspective on growth across OpenAI’s B2B portfolio. We connect monetization and customer economics with compute demand, capacity, margin, and resource planning, partnering closely with Product, GTM, Strategic Deal Desk, Compute Finance, Capacity Planning, Data, Accounting, and other Finance teams. About the Role We are hiring a Strategic Finance leader for B2B Product to build the 0→1 financial foundations and decision-making frameworks that will help OpenAI scale its B2B business sustainably. You will own high-impact work across B2B monetization, compute demand modeling, enterprise deal economics, contribution margin, and more. This is a portfolio-wide individual contributor role spanning our API Platform and other B2B products. You will connect customer demand and commercial terms to revenue, compute consumption, and margin outcomes; influence some of our largest enterprise deals; and help leadership make sound growth investment, compute capacity, and resource allocation decisions. This role is based in San Francisco, CA. In this role, you will: Own an integrated view of B2B monetization and product economics across the API Platform and other B2B products – including pricing, packaging, channel and usage mix, discounts, credits, commitments, revenue, and customer profitability Build, maintain, and improve complex driver-based financial models that connect customer usage, product and model mix, pricing, discounting, credits, commitments, and contribution margin Build financial views for B2B compute demand and margin planning; identify risks and opportunities, explain key drivers, and recommend capacity and resourcing decisions that improve portfolio economics Evaluate some of OpenAI’

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Network Engineering team within IT and Security advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient network services. We build and operate the connectivity that supports OpenAI’s offices, labs, campuses, cloud environments, people, and devices. By combining strong network fundamentals with security, reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at OpenAI, you will design, operate, and continuously improve the global networks that connect our offices, labs, campuses, PoPs, cloud environments, people, and devices. The role spans strategic platform engineering and responsive production operations: you will shape architecture, standards, roadmaps, lifecycle plans, and automation while supporting incidents, escalations, and time-sensitive delivery. Operational signals will inform what we stabilize, simplify, standardize, or automate next. We work backward from user needs, investigate root causes, own outcomes end-to-end, and move quickly without compromising security. We are looking for a versatile engineer who can make pragmatic reliability and security tradeoffs, communicate clearly, and turn recurring operational work into durable platforms, tooling, and standards. You will partner across IT, Security, AppEng, Research, Applied, workplace teams, carriers, and vendors. In this role, you will: Design, implement, and operate secure, scalable enterprise networks across offices, labs, campuses, PoPs, cloud connectivity, and hybrid environments. Set strategic direction for network services through architecture, standards, roadmaps, lifecycle planning, capacity strategy, and measurable reliability outcomes. Own production operations, including on-call, incident response, escalations, and time-sensitive delivery, while protecting user experience,

PythonAWSAzureCI/CD
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp

PythonAWSAzureRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Technical Support team is responsible for ensuring that developers and enterprises can reliably build mission critical solutions using OpenAI models. We provide technical guidance, resolve complex issues and support customers in maximizing value and adoption from deploying our highly-capable models. We work closely with Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. We think from an automation-first mindset and leverage the latest in AI to scale our support operations. Join the Senior Support Engineering (SSE) team at OpenAI and help shape the future of Technical Support in the age of AI. About the Role We are looking for a Senior Support Engineer to collaborate directly with our strategic enterprise accounts and product teams, helping solve some of the most difficult problems faced by our Customers. You will be part of the best technical troubleshooting team at OpenAI, and our Customers and Engineering teams will look to you for technical guidance in addressing the most technically difficult issues in our environment. As a Senior Support Engineer, you will design and run operational processes to monitor our top strategic customers and a 24x7 response team. You’ll work closely with our Infrastructure and Engineering teams to deliver the best possible experience to customers at scale. Working directly with our most strategic Customers - You will be crucial to the success of the most innovative, disruptive, and high-scale AI solutions being built with the OpenAI API platform. The nature of this role will be low volume, high difficulty. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Be among the foremost technical and troubleshooting experts for our API platform at OpenAI. You are the last line of defense before the core Engineering team. Proactively iden

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Fleet team builds core components to enable productive research from small to state of the art scale across OpenAI, with the goal of accelerating progress towards AGI. We frequently collaborate with other teams to speed up the development of new state-of-the-art capabilities. About the Role As we scale up with more researchers and engineers joining OpenAI, we seek a pragmatic and passionate engineer with a strong focus on the development experience for both engineers and scientists. In this role, you will be responsible for building and maintaining systems that allow our research + engineering organization to iteratively develop, test, and deploy new features reliably, with high velocity, and with a frictionless and fast development cycle. You will help oversee and drive to the vision of how we should build, test and deploy software. You will drive the design of our continuous integration pipelines, testing infrastructure, training and support around our build system. Our current environment relies heavily on Python, Rust, and C++, which you will take ownership of and strive to transform into a state of the art development experience for research. Ultimately, your role will be to provide the necessary tools and metrics to support our fast-paced culture and ensure a stable, scalable platform for growth, while also fostering a seamless and low friction experience for OpenAI’s research. This role is based in San Francisco, CA. For a San Francisco role, we use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if you: Have supported large monorepo development and deployment before Are a proficient Python programmer working in large monorepos Are proficient with Docker and Kubernetes Experienced in CI/CD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boun

PythonAWSDockerKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI's rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable infrastructure in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented and bring rigor to building and maintaining reliable systems. Demonstrate excellent software enginee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun

AWSAzureGCPKubernetes
V
📍 San Francisco, California, United States· Full-time
✓ High-confidence listing

$165K – $185K/yr

Quick readStrong listing-quality and freshness signals

About VSCO For years, we've helped photographers create their work. Now we're building what comes next. VSCO exists for photographers. Not as a side feature, not as an afterthought, but as the whole point. We build the connected system photographers rely on, and we've spent over a decade earning the trust of a global creative community that takes the craft seriously. Photography is at an inflection point. AI is reshaping what's possible for creative work, and that's where our mission shines. VSCO is building the full photographer's workflow: from creating and editing your work, to delivering it to clients, to running your business. All of it built thoughtfully, with photographers leading the way. We believe the future of photography tools creates more space for creativity, handling the busy work so photographers can focus on the craft. If you care about craft, community, and what technology can unlock for creative people, this is the work. We're a mission-driven and focused company where your work ships quickly, is meaningful, and reaches tens of millions of people worldwide. You'll have a real say in what we build and how we build it. We hire people who don't wait to be asked, naturally connect the dots, care about the quality of what they ship, and believe the best outcomes come from building together. About The Role We’re looking for a Senior Software Engineer, Infrastructure to own the platform that VSCO product and data teams ship on. You’ll join a small infra team that treats AWS, EKS, and GitOps as the default path for new systems, and you’ll spend real time pairing with other teams so - search, Workspace, and data - land in the right Terraform, Helm, and Flux instead of one-off snowflakes. This is a hands-on senior seat. You will design Terraform modules, cut production traffic to EKS, secure traffic with Cloudflare, look for optimizations in our AWS infrastructure, and leave on-call better than you found them. You like production ownership, you can sequence

PythonJavaSQLMySQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Product Marketing team shapes how customers understand, adopt, and realize value from OpenAI’s technology. We work across Product, Research, Sales, Solutions, Partnerships, and Customer Success to bring customer insight into our product strategy and translate technical capabilities into clear, credible stories and solutions. About the Role AI becomes meaningful when it helps people do the work that matters to them. For a finance team, that might mean understanding complex information faster. For a healthcare provider, it might mean navigating clinical workflows more effectively. For a sales or marketing team, it might mean creating entirely new ways to reach and serve customers. We’re looking for a senior product marketing leader to shape how OpenAI serves the business functions and industries where our technology can make a meaningful difference. You’ll define how our models and products meet the needs of teams such as sales, marketing, and finance, as well as industries including financial services, healthcare, and retail. Working closely with Product, Research, Sales, Solutions, and Partnerships, you’ll identify important customer problems, influence product strategy, and build relationships with the ecosystem partners and data providers needed to bring complete solutions to market. You’ll also build and lead the product marketing team responsible for turning these opportunities into durable customer value. You might thrive in this role if you: Have 12+ years of experience in product marketing, industry marketing, solutions marketing, or enterprise go-to-market, ideally across enterprise software, cloud, data, developer, or AI platforms. Have built and led high-performing teams, mentored senior marketers, and know how to create clarity in fast-moving, ambiguous environments. Understand how different industries and business functions evaluate technology, adopt new tools, and define value. Have shaped positioning and go-to-market strategies for c

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.5%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. AI and intelligent systems are driving the fifth paradigm shift, following previous technological revolutions like mainframes, personal computers, the internet, and mobile devices. We believe, in the foreseeable future, AI will revolutionize the FinTech industry - from how consumers understand and manage their finances, to how developers build applications and how all companies operate. The fintech industry landscape will undergo a fundamental reshape. Plaid in the FinTech AI Ecosystem Plaid is uniquely positioned to become the financial data and insights backbone for AI applications and platforms in this evolving ecosystem. We believe consumers should be able to understand and manage their financial life through conversational AI interfaces using natural language. We believe consumers should have peace of mind with a trustworthy consent and authorization manager when agents shop for them. We believe identity verification and financial fraud prevention in AI-powered products should feel seamless and embedded for the end users. The list goes on. The most important AI companies, major fintechs, and customer agent platforms are actively trying to integrate Plaid into AI-powered products and solutions t

JavaScriptJavaAWSMicroservices
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.5%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. AI and intelligent systems are driving the fifth paradigm shift, following previous technological revolutions like mainframes, personal computers, the internet, and mobile devices. We believe, in the foreseeable future, AI will revolutionize the FinTech industry - from how consumers understand and manage their finances, to how developers build applications and how all companies operate. The fintech industry landscape will undergo a fundamental reshape. Plaid in the FinTech AI Ecosystem Plaid is uniquely positioned to become the financial data and insights backbone for AI applications and platforms in this evolving ecosystem. We believe consumers should be able to understand and manage their financial life through conversational AI interfaces using natural language. We believe consumers should have peace of mind with a trustworthy consent and authorization manager when agents shop for them. We believe identity verification and financial fraud prevention in AI-powered products should feel seamless and embedded for the end users. The list goes on. The most important AI companies, major fintechs, and customer agent platforms are actively trying to integrate Plaid into AI-powered products and solutions t

AWSAIGoRust
🔔

Get new ai platform director jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime