Jobiba hiring network

Capacity Planning Lead Jobs

600 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current capacity planning lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Bazaarvoice
📍 Vilnius• Full-time• Hybrid
18 days ago

About Bazaarvoice At Bazaarvoice, we create smart shopping experiences. Through our expansive global network, product-passionate community & enterprise technology, we connect thousands of brands and retailers with billions of consumers. Our solutions enable brands to connect with consumers and collect valuable user-generated content, at an unprecedented scale. This content achieves global reach by leveraging our extensive and ever-expanding retail, social & search syndication network. And we make it easy for brands & retailers to gain valuable business insights from real-time consumer feedback with intuitive tools and dashboards. The result is smarter shopping: loyal customers, increased sales, and improved products. The problem we are trying to solve : Brands and retailers struggle to make real connections with consumers. It's a challenge to deliver trustworthy and inspiring content in the moments that matter most during the discovery and purchase cycle. The result? Time and money spent on content that doesn't attract new consumers, convert them, or earn their long-term loyalty. Our brand promise : closing the gap between brands and consumers. Founded in 2005, Bazaarvoice is headquartered in Austin, Texas with offices in North America, Europe, Asia and Australia. It’s official: Bazaarvoice is a Great Place to Work in the US , Australia, India, Lithuania, France, Germany and the UK! The Bazaarvoice services team is looking for candidates experienced with delivering Web technologies in a professional services capacity delivering the scope of work in an allocated number of hours and days. This individual will be responsible for all phases of the client’s projects including planning, design, implementation, configuration, data transformation, training and transition to other teams. The salary range starts from 2000 EUR gross per month (+additional language allowance) and offers are determined by experience, knowledge, and skillset. What You Will Do <div

javascriptSEO
View job →
O
OneTrust
📍 Bengaluru• Full-time
23 days ago

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We are looking for a Senior Project Manager who is passionate about delivering software implementations and building excellent relationships with our customers. As an Implementation Project Manager you will own client engagements and be the internal advocate for your customers. Your Mission You will be responsible for key project factors like resources, scope, timeline, deliverables, budget, reporting, and customer satisfaction. You will collaborate across organizations including Implementation, Sales, Product Management, and Support throughout the project lifecycle. As an Implementation Project Manager you will: Manage customer-facing professional service projects, responsible for tasks including requirements gathering, project planning, execution and status tracking, internal reporting and record keeping, customer reporting and stakeholder management Engage with new and existing enterprise customers in a project management capacity, working with other OneTrust Consultants to deliver the solution as scoped Build and manage customer-facing project plans and status repo

awsgitai
View job →
C
Cloudflare
📍 Remote• Full-time• Remote• $220K – $303K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations Remote US About the Role The Global GTM Strategy, Planning and Operations team at Cloudflare supports our high-growth Sales organization through scalable, AI supported GTM planning and operations. We focus on thoughtful planning of sales capacity, segmentation, territory carving, quota setting, and compensation design while ensuring clear rules of engagement for our Sales Team. Our goal is to establish a cohesive, consistent pr

REMOTEawsaigo
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Network Engineering team within IT and Security advances the mission of deploying artificial general intelligence (AGI) for the benefit of all by delivering secure, scalable, and resilient network services. We build and operate the connectivity that supports OpenAI’s offices, labs, campuses, cloud environments, people, and devices. By combining strong network fundamentals with security, reliability, automation, and user-centered design, we enable impactful AI research, corporate operations, and product innovation. About the Role As a Network Engineer at OpenAI, you will design, operate, and continuously improve the global networks that connect our offices, labs, campuses, PoPs, cloud environments, people, and devices. The role spans strategic platform engineering and responsive production operations: you will shape architecture, standards, roadmaps, lifecycle plans, and automation while supporting incidents, escalations, and time-sensitive delivery. Operational signals will inform what we stabilize, simplify, standardize, or automate next. We work backward from user needs, investigate root causes, own outcomes end-to-end, and move quickly without compromising security. We are looking for a versatile engineer who can make pragmatic reliability and security tradeoffs, communicate clearly, and turn recurring operational work into durable platforms, tooling, and standards. You will partner across IT, Security, AppEng, Research, Applied, workplace teams, carriers, and vendors. In this role, you will: Design, implement, and operate secure, scalable enterprise networks across offices, labs, campuses, PoPs, cloud connectivity, and hybrid environments. Set strategic direction for network services through architecture, standards, roadmaps, lifecycle planning, capacity strategy, and measurable reliability outcomes. Own production operations, including on-call, incident response, escalations, and time-sensitive delivery, while protecting user experience,

pythonawsazure
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and

REMOTEmachine learningaigo
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

REMOTEpythonawsazure
View job →
S
Supabase
📍 Remote• Full-time• Remote
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We are seeking a platform engineer to join our Compute Capacity team. This team owns the capacity plan that keeps compute supply ahead of demand across every region we operate in — the forecasting, buffer policy, and provisioning systems that make sure a project never runs into a wall it didn't know was there. You'll work on the systems that turn a capacity plan into provisioned reality: reservation acquisition, fleet reconciliation, and the automation that keeps what we've committed to in sync with what we're actually running. You'll help build the metrics and alerting that let capacity problems surface months out, on vendor lead time, rather than at the moment someone needs the room. You'll design, build, and operate systems that are both robust and highly automated — helping us hold the right buffer at the right cost, catch drift before it becomes a shortage, and give every team a single, trustworthy view of how much room we have across the millions of databases we manage. What You'll Be Responsible For Help build and maintain the capacity plan that keeps Supabase's compute supply ahead of demand across regions and instance families Support buffer policy by modeling headroom targets and their cost tradeoffs for review and sign-off Build and maintain automation that turns the capacity plan into provisioned reality — reservation acquisition and top-up, fleet reconciliation, drift detection between committed and running capacity Extend our infrastructure as code for capacity-relevant provisioning Instrument capacity: build and maintain metrics for saturation, reservation coverage, idle buffer, forecast error, and provisioning latency Build and tune capacity alerting so headroom,

REMOTEtypescriptpythonaws
View job →
S
Snowflake
📍 Bellevue• Full-time• Remote
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Capacity Engineering At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic, fast-moving environments and approach challenges with an experimental mindset, rapidly testing emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s infrastructure is expanding rapidly across AWS, Azure, and GCP. The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models demand, forecasts requirements, and delivers optimal CPU and GPU capacity on schedule. We drive hardware cost-efficiency and price/performance while continually maximizing fleet utilization. To achieve this across all major cloud providers, the team is building a centralized, self-se

REMOTEpythonjavaaws
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer on the Internal Tooling team, you will own the internal operating system that sits at the heart of how Baseten operates. Capacity helps unlock revenue by carefully balancing supply and demand. The operating system manages all aspects of the customer lifecycle: from onboarding to managing complex customer SLA requirements. This role is for engineers who want to own a product end to end, not just implement tickets. You will work directly with the Capacity, Sales, and Engineering teams to understand requirements, define solutions, and ship software that removes friction from some of the most high-stakes workflows in the company. If something is slow, manual, or error-prone in the capacity fulfillment lifecycle, you will be the one to fix it. You are a strong fit if you have strong product intuition, move fast without sacrificing quality, and take satisfaction in building tools that make the people around you measurably more effective. RESPONSIBILITIES Own the Capacity product end to end: scoping, design, implementation, and iteration based on feedback from internal stakeholders Translate complex operational requirements from Capacity, Sales, and SRE teams into clean, ergonomic product experiences Build and maintain full-stack features across the Capacity toolchain, including UI surfaces, APIs, and backend services Identify workflow bottlenecks and manual processes across the capacity lifecy

javascriptjavaaws
View job →
P
Pinterest
📍 San Francisco• Full-time• Remote
1mo ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Staff Software Engineer, Capacity Engineering. The team is responsible for efficiently managing one of the largest-scale cloud-native infrastructures in the world. This role is highly impactful, as efficiency is an ongoing strategic priority for Pinterest. The role has direct visibility across Pinterest Engineering and with Engineering and company leadership. The team is looking for a candidate with a strong background in implementing performance and efficiency projects on large scale distributed systems. In this individual-contributor role you will own and drive performance and efficiency for a core area of Capacity Engineering, partnering with the company-wide efficiency lead and collaborating with performance and efficiency leaders across the organization. What you’ll do: Drive efficiency in large-scale shared environme

REMOTEpythonjavaaws
View job →

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . What We're Looking For: 2–4 years of program or project management experience (or equivalent), ideally touching infrastructure, platform, or cloud environments. Working knowledge of cloud platforms (AWS or similar) — compute, storage, and basic usage/cost concepts. Interest in, and some exposure to, capacity management and infrastructure efficiency (right-sizing, utilization, waste reduction); deep FinOps expertise is not required. Strong organizational and execution skills: able to track a program's moving pieces, follow up, and keep things on schedule. Clear written and verbal communication; comfortable presenting status and asks to engineering partners. Collaborative and coachable — works well with engineers and more senior TPMs, seeks input, and takes feedback well. Experience at a large-scale consumer tech or infrastructure organization is

REMOTEsqlawsai
View job →
🔔

Get new capacity planning lead jobs by email

Daily job updates · Unsubscribe anytime