About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec
Jobiba hiring network
Capacity Planning Lead Jobs
600 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current capacity planning lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a member of the Capacity Strategy & Operations team, you will sit at the intersection of supply intelligence, demand forecasting, and cross-functional execution, turning a complex, fast-moving hardware market into a predictable, reliable foundation for our customers and internal engineering teams. This is not a purely analytical role. You will own the end-to-end capacity planning process: from translating customer commitments and growth forecasts into concrete supply requirements, to coordinating fulfillment across vendors, finance, and the infrastructure team, to building the systems that make all of this repeatable and scalable. When supply is constrained and tradeoffs are unavoidable, you are the person in the room who can model the options, make a clear recommendation, and drive alignment fast. You are a strong fit if you have operated at the intersection of strategy and execution before — someone who is equally comfortable building a capacity model in a spreadsheet and running a cross-functional war room when a customer deployment is at risk. EXAMPLE INITIATIVES Demand-Supply Alignment Framework: Build and own the process that translates customer pipeline, signed commitments, and growth projections into a forward-looking GPU demand signal — so the team is never caught flat-footed when a customer scales faster than expected. Constrained Allocation Playbook: Define the decision framework for how Basete
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Capacity & Efficiency Engineering team builds the software that manages, governs, and reduces Robinhood's AWS cloud spend, operating at the intersection of cloud infrastructure, data engineering, and FinOps. The team owns the full lifecycle of cloud cost: the data platforms that make spend transparent and attributable, the anomaly detection and forecasting systems that make it predictable, the automation that continuously rightsizes infrastructure at fleet scale, and the capacity planning and commitment strategy that keep a bursty, latency-sensitive trading platform both reliable and cost-effective. Our work carries CEO-level visibility and has already driven millions of dollars in annualized savings. We partner closely with Data Science, Infrastructure, Finance, and product
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe is building tools to help our users create and scale their businesses, and grow the GDP of the internet. The Workforce Management team at Stripe helps manage our growing support operations network of both internal and external resources by consulting on operational change initiatives, modeling support requirements, and managing workload and resource allocations for an optimal user experience. What you’ll do As a Strategic Capacity Planning Specialist, you will play a key role in designing and maintaining capacity planning models for one of our key user support or risk verticals, with a focus on user experience, budget, utilization, and speed of answer to our users' queries. You will operate in a fast-paced and energizing environment surrounded by top talent, where resilience and openness to feedback will be key to success. Responsibilities Develop and maintain capacity planning models to forecast resource requirements. Leverage forecasting techniques to analyze historical data and market trends, predicting future capacity needs and enabling informed planning decisions. Evaluate potential risks and opportunities regarding model inputs to inform planning decisions. Monitor performance against planned capacity levels and adjust plans as needed, while holding stakeholders accountable for the accuracy of plan inputs. Collaborate with the vendor delivery team and stakeholders to gather information and align planning efforts with Stripe'
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
About the Team The Stargate organization is responsible for building and scaling the physical infrastructure systems that power OpenAI’s next generation of AI training and inference platforms. This includes the manufacturing, deployment, and operational execution required to bring large-scale compute infrastructure online globally. The team operates at the intersection of data center infrastructure, hardware manufacturing, supply chain, deployment operations, and systems planning. We partner closely across Infrastructure Strategy, Manufacturing Operations, Capacity Planning, Supply Chain, Deployment, and Engineering to execute one of the largest infrastructure scale-outs in the industry. About the Role We are seeking a Technical Program Manager, Rack Delivery to drive operational execution across rack manufacturing, site readiness, and deployment coordination for Stargate infrastructure programs. This role will serve as a key connective layer between manufacturing partners, deployment teams, and infrastructure readiness programs to ensure rack production and delivery timelines remain aligned with site availability and deployment sequencing. You will help manage operational execution across contract manufacturers (CMs), support build planning and RCCA processes, and coordinate deployment readiness across multiple concurrent infrastructure programs. You will also partner closely with Demand Planning teams to translate strategic planning inputs into actionable SKU-level manufacturing and delivery schedules. This role is ideal for someone who thrives operating across ambiguity, manufacturing operations, infrastructure deployment, and large-scale cross-functional execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation support. Key Responsibilities Drive cross-functional coordination between rack manufacturing, deployment operations, and site readiness programs. Manage operational execution acros
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the production performance envelope around tokens, latency, reliability, cost, and capacity; the core execution loop and interfaces that turn models into useful behavior; the shared infrastructure that enables other teams to build on Codex; and the feedback loops that turn real-world usage into better models and better agent behavior over time. About the Role We’re looking for engineers to build the infrastructure that powers Codex agents in production. This role focuses on the systems that let models safely execute code, interact with tools, complete long-running tasks, and operate reliably and efficiently at scale. You’ll design and operate the infrastructure behind sandboxed execution, orchestration, stateful workflows, app-server and SDK boundaries, and model rollouts. You’ll work at the intersection of distributed systems, developer tooling, and AI, building primitives that make Codex faster, safer, more reliable, and easier for the rest of the organization to build on. What You’ll Do Design and build execution environments for AI agents, including sandboxing, isolation, and reproducibility. Develop systems for agent orchestration across multi-step, tool-using workflows. Build infrastructure for running, testing, and debugging code generated by models. Create state and memory systems that allow agents to persist context across long-running tasks. Optimize tokens, latency, reliability, and cost across Codex’s production fleet. Support model rollouts, capacity planning, and the core tradeoffs between quality, speed, and economics to manage a fleet of frontier agents at scale. Build shared platform capabilities that unblock product teams, partner teams, and open source Codex. Yo
ABOUT THIS ROLE GHX is standing up a new Human-in-the-Loop (HiTL) operations team in Hyderabad to support its ADM (Automated Document Management) platform — an AI-powered document processing system replacing legacy UiPath automation for GFax North America. The platform processes over 2.2 million Purchase Orders per year and targets a 75–85% fully-automated (touchless) rate. The team handles every document the AI cannot process autonomously. The Operations Supervisor is building this team from zero and is accountable for everything within it: throughput, field accuracy, team development, day-to-day operations management, and real-time coordination with GHX Engineering and Product teams in the United States. This is a high-visibility, inaugural role — the team’s accuracy and throughput directly determine the speed and confidence with which GHX can expand the ADM platform beyond GFax North America. KEY RESPONSIBILITIES People & Performance Management — 70% Directly supervise 15-20 Document Review Specialists; own rostering across rotational shifts, attendance tracking, and queue capacity planning against daily document volume Manage operations SLA – TAT, Quality, along with people metrics Own day-to-day operations management: shift rosters and week-offs, leave planning, planned & unplanned shrinkage, absenteeism control, and real-time reallocation of staffing to protect SLAs & operations performance metrics Set and communicate daily, weekly, and monthly throughput, AHT, and accuracy targets; monitor individual and team performance against SLA and productivity standards Participate in and present at weekly, monthly, and quarterly business reviews (WBR / MBR / QBR) — performance against KPIs, root-cause analysis of misses, and committed corrective actions Own the full performance cycle for direct reports, including goal setting, mid-year and annual appraisals, with documented performance ratings, development plans, and calibration against peer supervisors Con
JOB TITLE Cloud Compute Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. WHAT YOU'LL DO Design, build, and operate Kubernetes clusters on Amazon EKS, including cluster lifecycle management, networking, autoscaling, and workload scheduling. Manage and optimize EC2-based compute infrastructure, including instance selection, placement strategies, capacity planning, and utilization analysis. Operate and improve ECS-based services where applicable, ensuring consistency across our container runtime environments. Develop and maintain Infrastructure as Code (IaC) using Terraform to provision and manage compute resources at scale. Collaborate with development and platform teams to define compute patterns, containerization standards, and deployment best practices. Monitor compute environments for availability, performance, and cost, driving continuous optimization across the fleet. Contribute to architectural decisions around workload placement, multi-tenancy, OS image and container li
JOB TITLE Senior Storage Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO The Senior Storage Engineer will be responsible for monitoring, implementation, and management of the organization's File/Block storage infrastructure. This role requires to cover one of the 2 possible shifts in India as well as a deep understanding of File/Block storage technologies and cloud-based solutions, with a focus on ensuring data availability, scalability, and security. The ideal candidate will have extensive experience in storage engineering and a proven ability to manage complex storage environments. • Provide advanced troubleshooting and support for complex storage issues, minimizing downtime and ensuring seamless operations. • Oversee and optimize File/Block storage systems, ensuring high availability and performance. • Conduct capacity planning and forecasting to ensure adequate storage resources are available to meet future demands. • Monitor storage performance and work with US resource to find recommendations for improvements to enhance efficiency and reduce costs. • Implement integration and automation solutions and provide feedback for solution to US Team to streamline operations and improve storage management. • Implement robust security measures to protect data integrity and ensure compliance with industry regulations and standards. • Maintain comprehensive documentation of storage configurations and processes and provide training and guidance to junior team members. WHAT’S
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. Bridge, a Stripe company, is a rapidly growing business and the number of developers integrating our APIs is growing quickly. We provide Slack-based support to our developers on a wide range of topics in close partnership with our internal engineering team. We’re a small but mighty team that’s expanding global coverage and is looking to make our first Product Support Specialists hires in Mexico City. About the team In this role, you’ll be working directly with developers integrating Bridge APIs and helping them resolve their issues. You will take ownership of complex, technical user issues and work across teams to resolve them. As part of the team, you’ll have a big impact to grow the Product Support operation and enhance various aspects such as capacity planning and forecasting, operational tools and systems, workflow optimization and automation, metrics and reporting, quality control, and more. What you’ll do Stripe is launching Stripe Delivery Centers - a brand new global team to design, implement and grow Stripe’s operations for the next decade. We are looking for dynamic and curious people that have a passion for solving global user issues, building operations, driving process improvements and want to play a front-line role in building this new operational capability for Stripe and accelerating Stripe’s growth. If you like challenging, scaled problems and are an amazing teammate, we want to hear from you! Responsibilities Analyze and troubl
About the Team At OpenAI, Trust & Safety Operations is central to protecting OpenAI’s platform, customers, and the public from abuse. We partner closely with Product, Engineering, Legal, Policy and Go To Market teams to identify emerging risks, build and mature enforcement systems, and ensure high-integrity operations while delivering a great user experience at scale. We’re building the Monetization Trust & Safety Operations team to ensure OpenAI can grow advertising in a way that is safe, trusted, and sustainable—for users, advertisers, and the business. This team sits at the intersection of operational scale, product risk, and rapid revenue growth, designing systems and operations that enable ads to scale without compromising user trust or safety. It’s critical to us that our Ads product be built in a way that corresponds to our Ads principles , and this team is key to that. About the Role We’re looking for a senior operator with strong analytical instincts to help build and scale Monetization Trust & Safety Operations at OpenAI. In this role, you’ll flex across the team’s highest-priority data and operational needs—from reporting and dashboard insights to budget and capacity planning, project-based analysis, and data automation —while partnering closely with Product, Policy, Engineering, Legal, Go To Market, and Data Science and Data Engineering teams. This role sits at the intersection of strategy, execution, and data: you’ll define ambiguous problems, query and validate data, build decision-support systems, and translate operational signals into clear recommendations and scalable, AI-first solutions. You should be comfortable moving from a high-level question to a rigorous analysis, a useful dashboard, an automated workflow, or a durable operating mechanism. As OpenAI introduces new revenue-generating formats and partnerships, you’ll help the team understand where risks, capacity constraints, quality gaps, and opportunities are emerging. You’ll brin
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond to operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will design and develop systems that support the model lifecycle in production, including deployment orchestration, configuration management, operational automation, reliability, and capacity management. You will help transform complex operational processes into scalable platform capabilities that enable teams across OpenAI to deploy and manage models with greater confidence and less manual effort. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build and evolve the platform used to deploy, configure, and manage models across ChatGPT. Develop systems for deployment orchestration, model rollouts, operational visibility, and production readiness. Create abstractions and tooling that simplify complex infrastructure and improve the developer experience. Automate operational workflows, including incident detection, diagnosis, mitigation, and recovery. Improve the reliability, scalability, and efficiency of model deployments and the infrastructure that supports them. Build systems that support capacity planning, resource allocation, and infrastructure utilization. Partner with research, infrastructure, and product engineering teams to identify common chal
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 About This Role ClickUp is looking for a Senior Database Reliability Engineer to join our Database Operations team. You’ll be responsible for the performance, integrity, security, and availability of our PostgreSQL databases running on Linux in AWS. This role focuses on day-to-day database administration, operational excellence, and ensuring data is consistently reliable and well-managed across environments. Key Responsibilities Administer, monitor, and maintain PostgreSQL databases (250GB+) in production environments (AWS RDS, Aurora, EC2) Ensure database availability, performance, and data integrity through proactive monitoring and maintenance Perform routine database administration tasks including patching, upgrades, vacuuming, reindexing, and statistics management Execute and manage backup, restore, and disaster recovery procedures, including regular testing of recovery plans Handle user access management, roles, and database security to ensure compliance with best practices Perform capacity planning, storage management, and growth forecasting Troubleshoot database issues, including performance bottlenecks, locking/contention, and failed jobs Support application teams with query tuning, schema changes, and release deployments Manage database migrations, upgrades, and change requests with minimal downtime Maintain documentation for database configurations, standards, and operational procedures Participate in on-call rotations and provide support for production incidents Qualifications 7+ years of experience in a senior database administrator or Database Engineering role Strong hands-on PostgreSQL ad
Get new capacity planning lead jobs by email
Daily job updates · Unsubscribe anytime