ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use. As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently. This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVES Blog Posts https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/ https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/ https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid RESPONSIBILITIES Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference Work across the stack, from customer-facing features to low-le
Jobs in United States
Capacity Strategy And Operations in United States
286 active opportunities · Updated October 2026
Showing
15 jobs
Explore current capacity strategy and operations jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a People Business Partner to support our team during a critical phase of growth. This is a highly strategic, high-impact role for someone who has partnered closely with leadership teams and helped organizations scale with intention. You will be deeply embedded with managers, bringing clarity and rigor to how teams are structured, how leaders operate, and how talent is developed across the organization. You’ll help shape team effectiveness, identify critical talent gaps, drive talent and performance strategies that enable high-performing teams, and build the people practices and change management approaches that allow us to scale with both speed and discipline. This role requires strong business judgment and the ability to operate with deep context. You’ll partner closely with leaders to navigate complex organizational decisions, anticipate challenges before they surface, and bring a clear point of view on what great looks like at every level of the organization. RESPONSIBILITIES Strategic partnership to leadership Serve as the trusted people partner to leadership, maintaining deep business context and translating it into people priorities by proactively surfacing systemic issues, risks, and opportunities before they become urgent. Bring data-driven insights to advise management on org design, succession planning, performance, retention, and engagement. Build management capacity across the org, equ
$196.6K – $242.9K/yr
Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: The Director of Strategic Finance - GTM is the embedded finance partner to Drata’s GTM org, working directly with the CRO, CMO, SVP of CS, and SVP of BD. This role owns ARR forecasting, capacity planning, investment allocation and ROI. You will partner with the VP of RevOps and be a trusted advisor in the room with GTM leadership, not a downstream reporting function and significant
£215K – £260K/yr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role Cohere is seeking a Global Public Policy Manager to lead policy engagement on compute infrastructure, export controls, AI competitiveness, and sovereign AI strategies. Governments increasingly view AI as critical national infrastructure and are investing heavily in compute capacity, energy resources, and domestic AI ecosystems. This role will help position Cohere as a trusted partner in emerging discussions around AI infrastructure, national competitiveness, and sovereign AI deployment. Key Responsibilities Monitor and analyze developments related to AI infrastructure, data centers, energy policy, semiconductor policy, export controls, and national AI strategies. Develop policy positions on sovereign AI, compute access, digital sovereignty, and AI competitiveness. Support engagement with governments developing AI infrastructure investment programs and national AI initiatives. Collaborate with commercial, product, and corporate development teams on strategic opportunities involving public-private partnerships. Represent Cohere in policy discussions related to AI infrastructure, energy requirements, and technology c
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
From $10K/yr
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp's implementation and Advisory motion is evolving. As the business scales, we're building a System Integrator (SI) and Advisory channel that expands delivery capacity, reaches new customer segments, and drives faster time-to-value. This role sits at the center of that shift. This is a rare opportunity to build a channel program from the ground up at a company with the scale to make it matter. You'll have real ownership over how Ramp partners with System Integrators (SIs), Advisory firms, and boutique consulting shops. The goal is a repeatable, high-quality program that grows with Ramp and becomes a core part of how we acquire and deliver for customers. What You'll Do Acquire new System Integrators (SIs), Advisory firms, and boutique consulting shops as partners to grow the referral motion of partners brining new customers to Ramp Define how implementation work gets routed to System Integrator partners, building the qualification criteria and intake logic that makes it easy for sales team to bring the right partner into the right d
$200K – $240K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Platform org at Sentry is the engine that everything else runs on, spanning developer platform, core infrastructure underpinning the product, SRE, security, and IT. It's a broad, technically complex, and deeply consequential organization. We're looking for a Staff Technical Program Manager who can operate at the intersection of technical depth and strategic execution: someone who thrives in complexity, builds trust with senior engineering leaders, and has a gift for turning ambiguity into clarity and momentum. You'll report to the Head of Technical Program Management and work closely with the VP of Platform Engineering, their staff, and partner teams across Engineering, Product, and Design (EPD). This is a high-visibility role with real influence. You'll work directly with the CTO and senior leaders, shape how the Platform org operates, and help Sentry scale through one of its most important chapters. In this role, you will: Drive strategic execution. Partner with the VP of Platform Engineering and senior EPD leaders to translate priorities into clear, measurable plans. Own sequencing, milestones, and key decisions, and make sure leadership always has reliable visibility into progress and tradeoffs. Build data-driven delivery health. Establish the metrics and dashboards that reflect delivery confidence, risk, and engineering health across the Platform org. Keep planning and reporting high-signal and lightweight, focused on outcomes rather than activity. Manage capacity, dependencies, and risk. Create visibility into resourcing, cross-team dependencies, and constraints so leaders can align investment to the
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Role Description We are seeking a Strategic Finance Lead to serve as a key partner to Plaid’s Engineering and Product organization. This person will own product-level revenue and gross margin forecasting, as well as the analytical frameworks that guide how Plaid deploys capital and allocates resources across these areas. As the Strategic Finance Lead, Product & Engineering, you will lead the finance partnership for a set of critical product areas, working directly with the relevant Product and Engineering leaders. Responsibilities Product-Level Forecasting . Own revenue and gross margin forecasts for your product areas, balancing bottom-up builds with top-down sanity checks to arrive at the most robust point of view. Resource Allocation & Capacity Planning. Partner with Product and Engineering leaders on headcount planning, roadmap-based capacity forecasting, and OpEx discipline, surfacing and supporting the resolution of trade-offs. Performance Visibility . Own the recurring reporting cadence for your product areas, including KPI tracking, variance analysis, and the early identification of risks and opportunities. Cross-Functional Partnership . Serve as the primary finance partner to Produc
From $128K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor
From $1.1M/yr
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is seeking a dynamic Manager of Sales Development to lead, coach, and motivate our team of Sales Development Representatives (SDRs). In this high-impact role, you will directly influence decision-making and drive growth in a fast-paced environment. This is a US-based remote position reporting directly to the Director of Sales Development, with a preference for candidates located in the Central, Mountain, or Pacific time zones. If you are an innovative leader who challenges the status quo, asks sharp questions, and elevates team performance, we want to hear from you. Here is what we are looking for: You Will: Hire, train and manage a team of Sales Development Representatives Develop strategies for career growth within the Smartsheet sales organization Report on activity metrics and forecast to the sales executives/directors Motivate individuals and team to exceed objectives through coaching and mentorship Actively use Salesforce.com and other SaaS tools to manage sales process, and set standards for performance metrics You Have: 2+ years sales experience in SaaS 1-3 years sales management experience Passion for coaching and mentoring emerging sales talent Experience giving both positive and constructive feedback Demonstrated ability to collaborate with a distributed sales team Capability to understand customer pain points and requirements, capacity to respond with value of Smartsheet products and services Current US Perks & Benefits: Employer subsidized medical/vision and dental coverage f
From $1.1M/yr
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is seeking a dynamic Manager of Sales Development to lead, coach, and motivate our team of Sales Development Representatives (SDRs). In this high-impact role, you will directly influence decision-making and drive growth in a fast-paced environment. This is a US-based remote position reporting directly to the Director of Sales Development, with a preference for candidates located in the Central, Mountain, or Pacific time zones. If you are an innovative leader who challenges the status quo, asks sharp questions, and elevates team performance, we want to hear from you. Here is what we are looking for: You Will: Hire, train and manage a team of Sales Development Representatives Develop strategies for career growth within the Smartsheet sales organization Report on activity metrics and forecast to the sales executives/directors Motivate individuals and team to exceed objectives through coaching and mentorship Actively use Salesforce.com and other SaaS tools to manage sales process, and set standards for performance metrics You Have: 2+ years sales experience in SaaS 1-3 years sales management experience Passion for coaching and mentoring emerging sales talent Experience giving both positive and constructive feedback Demonstrated ability to collaborate with a distributed sales team Capability to understand customer pain points and requirements, capacity to respond with value of Smartsheet products and services Current US Perks & Benefits: Employer subsidized medical/vision and dental coverage f
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer leading Fleet Management, you will be the overall technical lead across three pods and the person who sets the technical direction for the fleet management layer of Roblox. This is a hands-on, deeply technical leadership role that owns all of Roblox's compute capacity end to end: from low-level provisioning and the data plane, up through the control planes that operate it, and all the way to the UI and internal-facing products that let teams self-serve capacity. Your org centralizes security, maintenance operations, and the uptime of every Roblox Kubernetes cluster, and governs the internal customer contracts that drive automation across the fleet spanning Roblox data centers and cloud providers. You will guide architecture, raise the engineering bar, and make sure compute capacity supply and demand stay in balance as the fleet grows. You will: Serve as the overall technical lead for three Fleet Management pods, setting and aligning the technical direction across low-level provisioning, the data plane, and the control plane and product surfaces above them. Architect the declarative, Kubernetes-style control planes that operate Roblox's compute fleet across o
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's Cache team is building a next-generation caching solution designed to deliver sub-millisecond average latency, horizontal scalability, and high efficiency—all at a drastically lower cost. Our ultimate vision is to shape a caching infrastructure capable of supporting 1 billion Daily Active Users while reducing costs by 90%. We are turning hours of onboarding and capacity expansion into seconds, freeing service owners entirely from managing cluster lifecycles. As a Senior Engineer on the Cache team (part of the Infra Storage org), you will innovate and operate large-scale, in-house distributed systems to solve Roblox's ever-growing caching challenges. You will report directly to the Engineering Manager for the Cache team. (Check out our recent engineering blog post here to learn more about the team's latest work!) You will: Lead the architectural transition to a next-generation, multitenant caching service built on ValKey, ensuring strict data, resource, and failure isolation for all tenants. Drive systemic optimizations to mitigate head-of-line blocking, manage hot keys, and maximize CPU and memory utilization across physical machine clusters. Design and build robust frameworks to a
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service discovery, secrets management and related software layers. We’re looking for a skilled Senior Site Reliability Engineer with strong programming skills to help us build Roblox's private cloud, productionize our growing Kubernetes-based infrastructure, and institute reliability best practices across the Roblox Compute team. You will: Design and Develop systems & libraries that promote fault-tolerance and resilience, automate much of the management and lifecycle of our clusters, and ensure systems are observable. Promote and Institute reliability best practices across the Infra Compute group, drive common reliability initiatives. Provides collaborative technical reviews and operational guidance to strengthen system reliability. Build, Automate and Standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem. Create Tooling that provides production guardrails, by evaluating release candidate capacity with load testing tooling before de
Other cities to consider
More places hiring for this role
Get new capacity strategy and operations jobs in United States by email
Daily job updates · Unsubscribe anytime