About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
Jobs in United States
Capacity Planning Lead in United States
293 active opportunities · Updated October 2026
Showing
15 jobs
Explore current capacity planning lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an
About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol
About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re hiring a Data Scientist to help build and scale our internal analytics capabilities. This is a foundational role where you’ll create dashboards, data models and insights to power business and product teams alike. You’ll collect requirements, define key metrics, and deliver insights directly to stakeholders. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape Baseten’s product and strategy. RESPONSIBILITIES Build and maintain production-grade dbt models and dashboards across multiple functions with a focus on accuracy, simplicity and user experience. Define and instrument core metrics around ROI, product adoption, customer lifecycle, capacity, availability, revenue and costs. Ingest and transform raw data using tools like dbt, Airbyte, and BigQuery. Partner with Engineering, Finance, Marketing, and Sales teams to understand goals and translate them into data solutions REQUIREMENTS 5+ years of experience in analytics engineering, data analysis, analytics, data science or a related role Advanced SQL and dbt skills, with a record of building models, tests, semantic layers and lineage in a cloud data warehouse. Prior experience supporting complex cross-functional projects across GTM, Finance and Engineering across various stages of the customer journey. Experience building dashboards and self-serve analy
About the Team OpenAI's Industrial Compute organization is building and scaling the infrastructure required to support frontier AI. The Infrastructure Strategic Sourcing team connects technical and project requirements to supplier readiness, contracting, purchasing, equipment delivery, and portfolio-level risk visibility across owner-furnished contractor-installed equipment (OFCI), data center networking, rack systems and integration, fiber, cabling, optical interconnects, and related infrastructure. The team partners across Pre-Construction, Design, Construction, Electrical and Mechanical Engineering, Network Engineering, Hardware and Rack Delivery, Strategic Sourcing, Procurement, Legal, Finance, Accounts Payable, Logistics, and external suppliers. We build the operating mechanisms that keep sourcing decisions, purchase execution, long-lead equipment, network and fiber dependencies, rack readiness, and delivery commitments aligned to infrastructure schedules. About the Role We are seeking an Infrastructure Sourcing Operations Lead to own procurement operations across pre-construction, design, construction, and sourcing through purchase order issuance, while maintaining visibility through invoice resolution, production, logistics, delivery, installation, and readiness. The portfolio includes electrical and mechanical OFCI, networking equipment, rack systems and integration, fiber, cabling, optical interconnects, and other infrastructure required to bring capacity online. In this role, you will set priorities, make or escalate decisions that affect cost, supplier relationships, contractual position, and delivery schedules, and define the standards used by execution support for queue management, documentation, tracker maintenance, and recurring reporting. Success requires sound commercial and program judgment, operational rigor, systems thinking, and the ability to turn incomplete information across vendors, tools, and project teams into clear decisions, accountable
About the Team At OpenAI, Trust & Safety Operations is central to protecting OpenAI’s platform, customers, and the public from abuse. We partner closely with Product, Engineering, Legal, Policy and Go To Market teams to identify emerging risks, build and mature enforcement systems, and ensure high-integrity operations while delivering a great user experience at scale. We’re building the Monetization Trust & Safety Operations team to ensure OpenAI can grow advertising in a way that is safe, trusted, and sustainable—for users, advertisers, and the business. This team sits at the intersection of operational scale, product risk, and rapid revenue growth, designing systems and operations that enable ads to scale without compromising user trust or safety. It’s critical to us that our Ads product be built in a way that corresponds to our Ads principles , and this team is key to that. About the Role We’re looking for a senior operator to help build and scale Monetization Trust & Safety Operations at OpenAI. In this role, you’ll own critical Ads T&S workstreams from problem framing through scaled operation, partnering closely with Product, Policy, Engineering, Legal, and Go To Market to turn ambiguous priorities into durable operating models. This role sits at the intersection of strategy and execution: you’ll define scope, align owners, manage milestones and risks, resolve dependencies, and build the workflows, decision structures, and operating mechanisms that allow Ads Trust & Safety to scale. You’ll move between standing up new programs, stabilizing existing workflows, and handing off durable ownership as priorities evolve. As OpenAI introduces new revenue-generating formats and partnerships, you’ll bring structure to complex initiatives that balance user safety, advertiser experience, and business growth. You’ll use data and frontline signals to identify bottlenecks, quality gaps, capacity needs, and high-leverage interventions, and communicate clear
From $194K/yr
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Coinbase has built the world's leading compliant cryptocurrency platform serving over 73 million accounts in more than 100 countries. With multiple successful products, and our vocal advocacy for blockchain technology, we have played a major part in mainstream awareness and adoption of cryptocurrency. We are proud to offer an entire suite of products that are helping build the crypto economy and increase economic freedom around the world. There are a few things we look for across all hires we make at Coinbase, regardless of role or team. First, we look for signals that a candidate will thrive in a culture like ours, where we default to trust, embrace feedback, disrupt ourselves, and expect sustained high performance because we play as a championship team. Second, we expect all employees to commit to our mission-focused approach to our work. Finally, we seek people with the desire and capacity to build and share expertise in the frontier technologies of crypto and blockchain, in whatever way is most relevant to their role. JOB DUTIES Scale and grow the HR Engineering team by hiring, onboarding and training new analysts and engineers to support Workday and HR functional area Provide functional & technical leadership and mentoring to team members Supervise the day-to-day activities of our Workday instance Design user-friendly processes, guidelines, and documentation
Work Flexibility: Onsite Join Stryker in a role that combines continuous improvement leadership with project portfolio management to support operational excellence across the site. This position leads cross-functional initiatives, applies Lean and Six Sigma methodologies, and manages strategic projects that align with business and manufacturing objectives. The role offers the opportunity to work across functions, drive process improvement activities, and support the advancement of a high-performing operational environment. What You Will Do Lead cross-functional continuous improvement and project management initiatives from scope definition through implementation, ensuring projects are delivered on time and within budget. Facilitate value stream mapping, Kaizen events, Gemba walks, and problem-solving activities to identify waste, improve process flow, and support productivity targets. Manage the site project portfolio, track progress against commitments, and communicate performance through standardized project management processes. Analyze operational and manufacturing data, conduct time studies, labor modeling, capacity assessments, and line balancing activities to support process optimization decisions. Coach leaders and teams on Lean manufacturing principles, Daily Management, Leader Standard Work, Six Sigma methodologies, and continuous improvement tools while maintaining the site continuous improvement knowledge base and 12-month Kaizen roadmap. What You Will Need Required: Bachelor’s degree in Business Administration, Industrial Engineering, or a related field. Minimum 5 years of experience leading continuous improvement, Lean manufacturing, operational excellence, or project management initiatives in a manufacturing enviro
Job Details: Job Description: As one of the world's largest semiconductor manufacturers, Intel is committed to advancing every aspect of semiconductor technology, from process development and manufacturing to advanced packaging and reliability characterization. Employees within Intel Foundry are part of a global network spanning technology development, manufacturing, assembly, test, and quality organizations across both front-end silicon and advanced packaging facilities. You will join the Foundry Lab Network (FLN), a key organization within Foundry Quality, Reliability, and Labs (FQRL), located at Intel's growing Chandler, Arizona site. This site serves as the technology development hub for Intel's most advanced packaging technologies, including Hybrid Bond Interconnect (HBI), Embedded Multi-die Interconnect Bridge (EMIB-T), glass substrates, Foveros, and future advanced packaging innovations. FLN is actively expanding its laboratory capabilities and capacity to support Intel Foundry's roadmap, enabling faster technology qualification and accelerating development cycles through Quick Turn Monitoring (QTM) solutions that significantly reduce time-to-data. As the Stress Test and Reliability (STaR) Back-End Laboratory Manager, you will play a critical leadership role within FLN. You will partner closely with Packaging Technology Development, Quality and Reliability, and Failure Analysis teams to establish and enhance laboratory capabilities that support qualification and certification of next-generation packaging technologies and products. This position offers a unique combination of people leadership, technical problem solving, and organizational strategy. You will lead a team of 7-10 talented engineers, spearhead cross-functional efforts to resolve complex reliability and technology challenges, drive strong quality and operati
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Fab Engineer in the RDA & Metrology Team , you will collaborate with other equipment engineers, technicians, and application owners. Your responsibilities include overseeing the installation, modification, upgrade, and maintenance of manufacturing equipment, providing technical support to manufacturing equipment repair and process engineering organizations. You will define and write preventative maintenance schedules, maintain records on technical notices, upgrades, and safety issues, and study equipment performance to establish solutions that improve tool uptime. Responsibilities: Drive performance standards, critical measures, and accountability for operational decisions and results. Lead equipment issue resolution, troubleshooting efforts, and root cause investigations. Manage project priorities and recommend changes, continuation, or cancellation to meet fab objectives. Collaborate daily with Equipment Leads to align priorities, resolve issues, and optimize resource utilization. Plan for capacity and product mix requirements while enhancing equipment flexibility and performance. Maintain and improve equipment reliability through maintenance procedures, best-known methods (BKMs), program monitoring, and documentation updates. Partner with vendors and multi-functional teams to implement continuous improvement initiatives that enhance throughput, productivity, and equipment health. Leverage AI, analytics, and
Job Details: Job Description: As the world's largest chip manufacturer, and a global leader in innovation and new technology, Intel strives to make every facet of semiconductor manufacturing state-of-the-art - from semiconductor process development and manufacturing, through yield improvement to packaging, final test and optimization, and world class Supply Chain and facilities support. Employees in the Technology Development and Manufacturing Group are part of a worldwide network of design, development, manufacturing, and assembly/test facilities, all focused on utilizing the power of Moore's Law to bring smart, connected devices to every person on Earth. We're constantly working on making a more connected and intelligent future, and we need your help. Change tomorrow. Start today. What we offer: We foster a collaborative, supportive, and exciting environment where the brightest minds in the world come together to achieve exceptional results. We give you opportunities to transform technology and create a better future, by delivering products that touch the lives of every person on earth. What we do: Join the Logic Technology Development (LTD) organization, a powerhouse of process technology innovation driving Intel's IDM 2.0 strategy to exceed the aspirations of our global customers. We are seeking a visionary Principal Engineer to lead the charge in Intel's foundry revolution, serving as a strategic technical architect and executive partner in guiding our most critical customers through the entire product development lifecycle. In this leadership capacity, you will orchestrate multidisciplinary strategies spanning device engineering, product integration, and yield analysis to steer customers from critical design tape-out and pre-silicon evaluation through to NPI silicon qualification. Furthermore, you will define t
Other cities to consider
More places hiring for this role
Get new capacity planning lead jobs in United States by email
Daily job updates · Unsubscribe anytime