Jobiba hiring network

Capacity Planning Lead Jobs

600 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current capacity planning lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

R
Remote
📍 Singapore• Full-time• Remote• From $46.5K/yr
22 days ago

About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage international teams. With our core values at heart and future focused work culture, our team works tirelessly on ambitious problems, asynchronously, around the world. You can find Remoters working from 6 different continents (Antarctica left to go!) and all of our positions are fully remote. With Innovation as one of the core values, we have built Automation and AI capabilities into the requirements for every role. We encourage every member of the Remote team to bring their talents, experiences and culture to the table to help us build the best-in-class HR platform. If you are energetic, curious, motivated and ambitious, be part of our world. Apply now and define the future of work! What this job can offer you As a critical extension of our clients’ HR teams, we navigate and advise on the intricate paths of global employment with unmatched speed, expertise, and precision. The vision of the Lifecycle team at Remote is not just about maintaining the gold standard in HR practices; it’s about elevating it, integrating cutting-edge technology solutions, and enriching customer experiences in over 80+ countries. On this team, your work directly influences our ability to sustain and extend our compliance coverage, continuously enhance our customer journeys, and significantly increase our operational capacity. You’re not just part of a team; you’re at the forefront of shaping the future of work, ensuring every interaction is fast, intuitive, and profoundly impactful. Dive into a role where your passion for innovation, commitment to excellence, and drive to make a global difference aligns with our mission to empower organizations worldwide to employ anyone, anywhere — compliantly. The Senior Lifecycle Specialist, Employee Relations & Transitions independently manages

REMOTEawsrestai
View job →
R
Remote
📍 Japan• Full-time• Remote• From $46.5K/yr
22 days ago

About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage international teams. With our core values at heart and future focused work culture, our team works tirelessly on ambitious problems, asynchronously, around the world. You can find Remoters working from 6 different continents (Antarctica left to go!) and all of our positions are fully remote. With Innovation as one of the core values, we have built Automation and AI capabilities into the requirements for every role. We encourage every member of the Remote team to bring their talents, experiences and culture to the table to help us build the best-in-class HR platform. If you are energetic, curious, motivated and ambitious, be part of our world. Apply now and define the future of work! What this job can offer you As a critical extension of our clients’ HR teams, we navigate and advise on the intricate paths of global employment with unmatched speed, expertise, and precision. The vision of the Lifecycle team at Remote is not just about maintaining the gold standard in HR practices; it’s about elevating it, integrating cutting-edge technology solutions, and enriching customer experiences in over 80+ countries. On this team, your work directly influences our ability to sustain and extend our compliance coverage, continuously enhance our customer journeys, and significantly increase our operational capacity. You’re not just part of a team; you’re at the forefront of shaping the future of work, ensuring every interaction is fast, intuitive, and profoundly impactful. Dive into a role where your passion for innovation, commitment to excellence, and drive to make a global difference aligns with our mission to empower organizations worldwide to employ anyone, anywhere — compliantly. The Senior Lifecycle Specialist, Employee Relations & Transitions independently manages

REMOTEawsrestai
View job →
R
Remote
📍 Singapore• Full-time• Remote• From $46.5K/yr
22 days ago

About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage international teams. With our core values at heart and future focused work culture, our team works tirelessly on ambitious problems, asynchronously, around the world. You can find Remoters working from 6 different continents (Antarctica left to go!) and all of our positions are fully remote. With Innovation as one of the core values, we have built Automation and AI capabilities into the requirements for every role. We encourage every member of the Remote team to bring their talents, experiences and culture to the table to help us build the best-in-class HR platform. If you are energetic, curious, motivated and ambitious, be part of our world. Apply now and define the future of work! What this job can offer you As a critical extension of our clients’ HR teams, we navigate and advise on the intricate paths of global employment with unmatched speed, expertise, and precision. The vision of the Lifecycle team at Remote is not just about maintaining the gold standard in HR practices; it’s about elevating it, integrating cutting-edge technology solutions, and enriching customer experiences in over 80+ countries. On this team, your work directly influences our ability to sustain and extend our compliance coverage, continuously enhance our customer journeys, and significantly increase our operational capacity. You’re not just part of a team; you’re at the forefront of shaping the future of work, ensuring every interaction is fast, intuitive, and profoundly impactful. Dive into a role where your passion for innovation, commitment to excellence, and drive to make a global difference aligns with our mission to empower organizations worldwide to employ anyone, anywhere — compliantly. The Senior Lifecycle Specialist, Employee Relations & Transitions independently manages

REMOTEawsrestai
View job →
I
Instawork
📍 Bengaluru• Full-time• From ₹3.3L/yr
22 days ago

Instawork is on a mission to create meaningful economic opportunities for skilled hourly professionals in communities around the globe. Our AI-powered labor marketplace helps local businesses scale, and enables global technology companies to push the frontiers of robotics and AI. Backed by world-class investors like Benchmark, Spark Capital, Craft Ventures, Greylock, Y Combinator, and others, we’re looking for exceptional talent to reimagine the way the world works. About the Role We are looking for experienced Shift Supervisors to oversee evening and night shift operations across our Airbnb, kitchen, and in-office data collection locations in Marathalli. As a Supervisor, you will manage a team of Data Collectors, ensure shift productivity and quality targets are met, and submit all required reports on time. Your specific shift timing and location will be confirmed during the interview process based on current requirements. Key Responsibilities • Oversee shift operations for a team of Data Collectors at the assigned location • Monitor productivity and quality in real time throughout the shift • Collect the next day's task list from the client before the stipulated time and coordinate with the Team Lead • Ensure demo recordings are submitted on time per the daily schedule • Flag quality issues immediately to the Trainer and Team Lead • Maintain shift attendance log and real-time productivity tracking • Submit SOD, MOD, EOD, and Inventory reports for the shift • Coordinate with the Inventory Manager on device availability and equipment issues • Support replacement of underperforming Data Collectors in coordination with the Team Lead What We're Looking For • Minimum 1 year of experience in floor management, team handling, or operations supervision • Graduate preferred — 10+2 minimum • Not a fresher — prior experience in a supervisory or team lead capacity is mandatory • Comfortable working for morning, afternoon or night shifts • Strong attention to detail and ability

gitrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
23 days ago

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

awskuberneteslinux
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

REMOTEkubernetesaigo
View job →
M
Mongodb
📍 Japan• Full-time
25 days ago

MongoDB Pre-Sales Solutions Architects are technical business advisors who help customers design, justify, and adopt reliable, scalable systems using MongoDB’s data platform. They own the technical strategy across complex opportunities—from discovery and qualification through architecture, proof of value, executive alignment, and successful adoption—and connect technical decisions to measurable business outcomes. You’ll partner closely with Account Executives, Sales Leadership, Customer Success, Professional Services, and ecosystem partners to shape multi-threaded account strategies, build champions, de-risk complex architectures, and drive expansion. You’ll serve as a trusted advisor to developers, architects, operations leaders, and business executives, helping organizations modernize legacy systems, build AI-powered applications, and realize measurable value from MongoDB. . We are looking to speak to candidates who are based in Tokyo for our hybrid working model. As an ideal candidate, you will have: Ideally 8 to 11 years of related experience in a customer facing role, with 5 to 7 years of experience in pre-sales with enterprise software Minimum of 3 years experience with modern scripting languages (e.g. Python, Node.js, SQL) and/or popular programming languages (e.g. C/C++, Java, C#) in a professional capacity Experience designing with scalable and highly available distributed systems in the cloud and on-prem Demonstrated ability to lead architecture reviews for complex, multi-component applications and platforms, identifying risks, evaluating trade-offs, and providing clear guidance to modernize, optimize, and de-risk the solution Excellent presentation, communication, and interpersonal skills, with the ability to convey complex technical and business concepts in a clear and compelling manner to technology and business leadership Ability to partner with Sales Leadership and Account Executives on multi-threaded account and territory strategies, prioritize oppor

pythonjavanode.js
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
26 days ago

About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal

REMOTEawsrestai
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
29 days ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it

REMOTEmachine learningaigo
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
29 days ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an

REMOTEmachine learningaigo
View job →
O
1mo ago

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to develop and operate increasingly capable AI systems at global scale. The Manufacturing Operations team works across Hardware Engineering, Manufacturing Engineering, Manufacturing Quality, Rack Integration, System Enablement, Supply Chain, Logistics, Deployment, and Hardware Operations to convert complex hardware designs into reliable production systems. We partner closely with ODMs, JDMs, contract manufacturers, and component suppliers to ensure that servers, racks, networking equipment, and supporting infrastructure are manufactured, validated, and delivered at the quality and scale required by OpenAI. About the Role We are seeking a Technical Program Manager to lead manufacturing operations programs across OpenAI’s hardware supply base in Singapore and the broader APAC region. You will own cross-functional execution from new product introduction through production ramp and sustaining operations. You will coordinate manufacturing partners and internal engineering teams around factory readiness, capacity, material availability, build plans, validation, quality gates, issue resolution, and delivery commitments. This role requires strong technical fluency, disciplined program management, and the ability to operate directly with manufacturing partners in fast-moving, high-stakes environments. You should be comfortable working at both the factory floor and executive-review levels, translating complex manufacturing conditions into clear risks, decisions, and recovery plans. This role is based in Singapore and requires regular travel to manufacturing partners across the APAC region. Key Responsibilities Lead manufacturing operations programs for AI servers, racks, networking systems, and related infrastructure across regional manufacturing partners. Own integrated program plans spanning NPI, factory readiness, material availability, tooling, test development, qualification,

REMOTEpythonsqlaws
View job →

About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol

awsrestai
View job →

About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.

awsazurerest
View job →
O
Okta
📍 Bengaluru• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. With the Okta's Auth0 organization’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical Reporting to the Manager of Engineering, in this role as a SRE Operations Engineer, you will ensure smooth operations of our Customer Identity Cloud at Okta. Working closely with the SRE team, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term operational success for the platform with potential career growth into Site Reliability Engineering. What you’ll be doing Executes operational work including updating/patching and maintaining the Engineering Service Desk queue Responsible for ensuring team requests are triaged and/or actioned in a timely manner Monitors Platform health and take steps to alleviate issues related to deployment and operations Assist with capacity, performance and scalability testing where required Escalation point for Platform issues from customer support teams Execute runbooks and update processes as required Interface with the SRE team to report core issues, required improvements and new feature requests What you’ll bring to the role General platform infrastructure knowledge, including high availability / l

nodejsmongodbaws
View job →
M
Mongodb
📍 London• Full-time
1mo ago

We are looking for passionate technologists to join our Pre-Sales organization to ensure that our growth is grounded and guided by strong technical alignment with our platform and the needs of our customers. MongoDB Pre-Sales Solution Architects are responsible for guiding our customers and users to design and build reliable, scalable systems using our data platform. Our team is made up of seasoned technical sales professionals, software architects, entrepreneurs, and developers who take direct responsibility for customer success, including the design of their software, deployment, and operations. You'll work closely with our sales executives, helping customers solve business problems by leveraging our solutions, playing a key role in winning deals and driving the business forward. You'll be a trusted advisor to a wide range of users from startups to the world's largest enterprise IT organizations. We are looking to speak to candidates who are based in London for our hybrid working model. As an ideal candidate, you will have: Ideally 8 to 11 years of related experience in a customer facing role, with 5 to 7 years of experience in pre-sales with enterprise software Minimum of 3 years experience with modern scripting languages (e.g. Python, Node.js, SQL) and/or popular programming languages (e.g. C/C++, Java, C#) in a professional capacity Experience designing with scalable and highly available distributed systems in the cloud and on-prem Demonstrated ability to work with customers to review complex architecture of existing applications, providing guidance on how to improve by leveraging technology Excellent presentation, communication, and interpersonal skills, with the ability to convey complex technical and business concepts in a clear and compelling manner to technology and business leadership Ability to strategize with sales teams and provide recommendations on how to drive a multi-threaded account strategy, aligning other MongoDB and ecosystem resources to

pythonjavanode.js
View job →
🔔

Get new capacity planning lead jobs by email

Daily job updates · Unsubscribe anytime