ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b
Jobiba hiring network
Capacity Planning Lead Jobs
600 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current capacity planning lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le
AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About the Role We're seeking a Revenue Operations Manager with a strong track record, a builder's mindset, and a bias for action to join our in-person team in New York or SF. This is a high-impact, hands-on role. You'll own the entire revenue operations function, from top-of-funnel lead routing through deal close and commission administration. You'll work closely with our Head of Finance & People Ops and sales leadership to build the systems, dashboards, and processes that scale our go-to-market motion. What You'll Do: Own the lead routing process from inbound and partnering with marketing to ensure proper attribution Run effective territory management & strategy for Geo based decisioning Support & strategise every aspect of revenue operations in your territory Own the strategy for capacity forecasting, inputs, throughputs & outputs being the conduit back to finance in
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We're looking for a Procurement Lead to build Baseten's procurement function from the ground up. This is a foundational, high-ownership role for someone who wants to define how a company runs procurement, not inherit an existing process. As the company scales across engineering, GTM, and corporate functions, we need to build purchasing workflows and vendor relationships. This is a rare, critical-moment hire: you'll set the leading-practice processes, systems, and policies that the company runs on as it continues to grow, with the mandate and dedicated focus to build something built to scale. This role exists to make procurement a source of leverage for the business, not friction. Your goal is to build scalable process, take on vendor management and RFP work that today sits informally with individual teams, and free up business owners' time and capacity, all while building the guardrails the company needs as it grows. You'll partner closely with Finance, Legal, Security, Compliance, IT, and business owners across the company, and you'll be judged on whether teams feel supported and unblocked, not slowed down. RESPONSIBILITIES Procurement Strategy and Process Design and implement Baseten's first company-wide procurement process, approval routing, and other workflows Establish procurement policies (spend thresholds, approval authority, competitive bidding requirements) tailored to a high-growth business Design pr
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
About Us What if your work could drive change in a globally established industry, shaping processes that touch every corner of the world? At Forto, we are at the forefront of change, harnessing the power of AI to revolutionise logistics. We want to reinvent digital supply chains to be transparent, frictionless and sustainable. From day one, our mission has been to simplify global trade – creating a seamless and efficient logistics process. Your role & Mission At Forto, we are made to move forward. We are transforming one of the world’s most complex industries by solving what others still treat as unsolvable. As Team Lead Product Manager for Financials, you will lead the product team responsible for the systems that turn commercial commitments, costs, revenue, and margin into reliable financial execution and business control. You will own the Financials product domain - one of Forto’s most critical areas - ensuring accurate revenue capture, cost control, margin protection, and trusted financial visibility. This is not a traditional finance tooling role. It sits at the intersection of logistics and finance, where complex operational challenges, partner costs, customer agreements, and fragmented processes must be transformed into scalable product solutions. You will also lead and coach two Product Managers working on closely connected domains, shaping the team's direction, raising the quality of product thinking, and helping them grow into stronger product leaders. This role is for someone who thrives in complexity, thinks at both system and detail level, influences across functions, and turns ambiguity into clear product direction. What you will do Lead the PM team responsible for Forto’s financial execution flow: Financials, procured services, committed capacity, cost structures, and margin control. Own the Financials product strategy across revenue, cost, invoicing, VAT, FX, reporting, financial controls, profit visibility, and customer-facing financial flows. T
At Secureframe , we are not just a company; we are at the forefront of revolutionizing cybersecurity compliance. Recognized as one of the industry's most innovative and trusted providers, Secureframe has consistently received accolades for our advanced technology solutions and commitment to excellence. With a robust portfolio of products that safeguard thousands of businesses worldwide, we have been featured in major publications such as Forbes’ next billion dollar startups , TechCrunch, and The Wall Street Journal for our transformative impact on the way companies achieve and maintain compliance standards. As we continue to grow, our mission remains clear: to provide seamless, secure solutions that enable businesses to focus on what they do best. Joining Secureframe means becoming part of a team dedicated to professional excellence and continuous learning in an environment that values creativity and forward-thinking. Secureframe is backed by top VCs including Kleiner Perkins, Accomplice, Gradient Ventures (Google’s AI Fund), BoxGroup, Village Global, and many more. As a member of our product design team, you'll take ownership of a Secureframe product area. You'll be involved in the entire product development process – from early research and visioning to adjusting pixels to launching. You'll define and refine our design system, as well as influence team culture and how we work cross-functionally together. Help iterate on our design processes, from design critiques to product workshops. We're a distributed, remote-first company (and have company off-sites when possible) that cares about collaboration, ownership, and growth. If you enjoy complex problems, improving customer experiences, and iterating with a collaborative team, we’d love to hear from you! Location & Work Model While Secureframe is a remote-first company, for this role, we have a strong preference for candidates who can work in a hybrid capacity from NYC. Remote candidates are still encouraged to a
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted. This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority. What You'll Own Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
Supabase is the open-source Postgres development platform that 7M+ developers and thousands of enterprises depend on every day. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the role We’re hiring a Product Manager to own the platform primitives every Supabase product runs on, including internal features like compute , disks , networking, and the API gateway, and external features like read replicas , custom domains , PrivateLink , and Bring Your Own Cloud. What you'll be responsible for: Talk to customers across the full spectrum. Indie developers running a single nano project, fast-growing startups whose costs are dominated by compute and disk, enterprises walking through a network-architecture review, and partners building on top of Supabase. Find the real blockers and bring them back to the roadmap. Own the problem statement and requirements behind every platform bet. Capture the customer evidence behind each decision, name the cost, capacity, and reliability constraints, and give the team a target it can hit. Decide what gets built, what gets deferred, and what gets cut. Every quarter you're choosing between enterprise unlocks blocking deals, reliability and cost wins for the long tail of projects, and net-new capabilities that change what Supabase can run. Set the priorities and defend them. Define how each launch is measured before it ships. Set the metric, agree on the threshold, and track it after launch. Know whether a feature moved enterprise deal velocity, project economics, or platform reliability. Use that to sharpen the next call. Keep engineering, design, and leadership aligned. The platform touches every other Supabase product, every region, and every customer tier. Write the roadmap, surface dependencies before they become blockers, and keep decisions moving. You might be a good fit if you: Have 7+ years of produ
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role Cohere is seeking a Global Public Policy Manager to lead policy engagement on compute infrastructure, export controls, AI competitiveness, and sovereign AI strategies. Governments increasingly view AI as critical national infrastructure and are investing heavily in compute capacity, energy resources, and domestic AI ecosystems. This role will help position Cohere as a trusted partner in emerging discussions around AI infrastructure, national competitiveness, and sovereign AI deployment. Key Responsibilities Monitor and analyze developments related to AI infrastructure, data centers, energy policy, semiconductor policy, export controls, and national AI strategies. Develop policy positions on sovereign AI, compute access, digital sovereignty, and AI competitiveness. Support engagement with governments developing AI infrastructure investment programs and national AI initiatives. Collaborate with commercial, product, and corporate development teams on strategic opportunities involving public-private partnerships. Represent Cohere in policy discussions related to AI infrastructure, energy requirements, and technology c
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! ========================== 📌 This is an active role, but we're currently building our talent pipeline so we're ready to move quickly once we have capacity to begin interviewing. At this stage, we're reviewing and shortlisting applications rather than scheduling interviews, so we aren't able to provide individual timelines or application status updates. If you're selected to move forward, we'll reach out as soon as we're ready to begin the interview process. To help us keep the process fair and efficient, we kindly ask that you don't follow up with the hiring team directly, as we won't have any additional updates to share until interviews begin. Thank you for your patience—we're looking forward to connecting with you when the process gets underway. ========================== Why This Role? We're hiring our first Senior Manager, Talent Development to help define how Cohere develops managers, leaders, and talent as we scale. This is a rare opportunity to own and shape a function from the ground up. Rather than inheriting a mature learning organization, you'll design the programs, frameworks, and ways of working that help one of th
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp's implementation and Advisory motion is evolving. As the business scales, we're building a System Integrator (SI) and Advisory channel that expands delivery capacity, reaches new customer segments, and drives faster time-to-value. This role sits at the center of that shift. This is a rare opportunity to build a channel program from the ground up at a company with the scale to make it matter. You'll have real ownership over how Ramp partners with System Integrators (SIs), Advisory firms, and boutique consulting shops. The goal is a repeatable, high-quality program that grows with Ramp and becomes a core part of how we acquire and deliver for customers. What You'll Do Acquire new System Integrators (SIs), Advisory firms, and boutique consulting shops as partners to grow the referral motion of partners brining new customers to Ramp Define how implementation work gets routed to System Integrator partners, building the qualification criteria and intake logic that makes it easy for sales team to bring the right partner into the right d
At Bazaarvoice, we create smart shopping experiences. Through our expansive global network, product-passionate community & enterprise technology, we connect thousands of brands and retailers with billions of consumers. Our solutions enable brands to connect with consumers and collect valuable user-generated content, at an unprecedented scale. This content achieves global reach by leveraging our extensive and ever-expanding retail, social & search syndication network. And we make it easy for brands & retailers to gain valuable business insights from real-time consumer feedback with intuitive tools and dashboards. The result is smarter shopping: loyal customers, increased sales, and improved products. The problem we are trying to solve : Brands and retailers struggle to make real connections with consumers. It's a challenge to deliver trustworthy and inspiring content in the moments that matter most during the discovery and purchase cycle. The result? Time and money spent on content that doesn't attract new consumers, convert them, or earn their long-term loyalty. Our brand promise : closing the gap between brands and consumers. Founded in 2005, Bazaarvoice is headquartered in Austin, Texas with offices in North America, Europe, Asia and Australia. It’s official: Bazaarvoice is a Great Place to Work in the US , Australia, India, Lithuania, France, Germany and the UK! The Bazaarvoice Content & Creators (C&C) Technical Services team is hiring an Implementation Engineer with experience delivering custom integration technologies in a professional services capacity. As a Technical Services Engineer, you will partner closely with client-facing and engineering colleagues in the US, UK, and Australia to deliver and support Bazaarvoice C&C products and services. This is a hybrid role that incorporates both business and technical responsibilities.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role The Palantir platform is deployed in numerous critical mission environments including combat zones and classified networks—from the back of a Humvee to a command post to the cloud. This means operating in multiple cloud environments, on-prem air-gapped networks, and at the edge—at scale. We are looking for Edge Infrastructure Engineers to build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure. This role demands a deep focus on low-level systems, including the deployment and management of physical bare metal servers in both traditional data centers and edge environments. You will be responsible for physical network engineering and the development of robust infrastructure that ensures performance of the Palantir platform. In addition to ensuring performance and reliability, you will play a critical role in building and scaling new environments in a forward-deployed capacity, including onsite. Edge Infrastructure Engineers combine hardware-level engineering experience with the drive to improve existing systems and the creativity to develop novel solutions for evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate for and participate in sensible, scalable systems design, sharing responsibility for diagnosing, resolving, and preventing production issues across our most demanding deployments. Core Responsibilities Maintaining availability of cloud & physical Ku
Get new capacity planning lead jobs by email
Daily job updates · Unsubscribe anytime