Jobs in United States

Capacity Strategy And Operations in United States

286 active opportunities · Updated October 2026

Explore current capacity strategy and operations jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Orchestration pod within Engine Productivity, you'll design and run the platform that executes large-scale end-to-end and integration tests, running the real, shipping client on real hardware, across Roblox's data centers, cloud, and our own device labs, so our engineering teams can ship the engine, clients, Studio, and more with speed and confidence. Every Roblox engine, client, and Studio change, along with the experiences built on top of them, should ship with confidence, and the Orchestration team is the layer that makes that possible. We build large-scale distributed services that turn thousands of test suites into a reliable, push-button pipeline: fanning work out across fleets of machines and real devices, moving artifacts to where they're needed, managing single- and multi-client test state, and giving test owners and maintainers a system to validate their own runs. It looks a lot like building a specialized cloud platform, with capacity-aware scheduling, isolation and sandboxing, and smart retry and backoff, plus the classic distributed systems problems (fairness, efficiency, failure handling, and reliability) at Roblox scale. You Will: Design a

PythonAWSGitAI
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $164K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: We are looking for an Advanced Analytics Lead to help Airbnb enable travel for our millions of guests and hosts on our platform. This role will sit under the Advanced Analytics family and support Product and Business leaders within our CS organization. The Difference You Will Make: Data thought partner to product and business leaders across teams through providing insights, recommendations, and enabling data informed decisions Drive day to day analytics and create scalable data tools Identify pain points in customer support operations and work with product leadership to improve experiences for our guest, host and agent community In addition, you will leverage Airbnb’s rich and unique data, state-of-art machine learning infrastructure, and other central data science tools to build and grow the measurement capacity within the organization. You will also be deeply involved in the technical details of the various systems we build, and will have the opportunity to collaborate with a strong team of engineers, product managers, designers and operations agents to achieve shared, cross-functional goals to help keep Airbnb’s community safe and trusted. You are passionate about solving complex problems within the Community support domain with data & insights; adding a new perspective to existing solutions and making business decisions based on careful and thoughtful analysis You are highly proficient in building and analyzing analytical frameworks, statistical models and experimentation methods to establish and communicate causal relationships You are a story

PythonSQLMachine LearningAI
D
📍 United States· Full-time· Remote
✓ High-confidence listing

$220K – $275K/yr

Quick readStrong listing-quality and freshness signals

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. Discord exists to give people the power to create space to find belonging — to talk regularly with the people they care about and build genuine relationships with friends and communities close to home or around the world. We're looking for an Analytics Manager to lead our Scaled Abuse Countermeasures and Research (SCAR) team — the team that safeguards Discord’s platform integrity. SCAR detects, analyzes, and disrupts high-volume threats through a combination of automated systems, deep research, and active incident response. This role reports to the Head of Safety Intelligence and Automation. What You'll Be Doing Lead and mentor a team of data analysts, scientists, and researchers who investigate active threats, identify platform abuse vectors, and uncover adversarial patterns. Define a strategic roadmap that prioritizes and disrupts the highest impact abuse operations through structured research, rigorous analyses, and live experimentation. Collaborate closely with the safety machine learning team to improve models by identifying the threat signals that translate into long-term, automated countermeasures. Partner cross-functionally with Product, Engineering, Data Science, Policy, Legal, and Revenue, influencing safety-by-design decisions upstream of abuse. Influence capacity toward high-impact infrastructure, such as automated rule engines, ML models, or agent moderation tools. What you should have 2+ years of people management experience leading technical teams, including engineers, data scientists, analysts, applied researchers, or equivalent. 4+ years of experience working in a Trust & Safety dom

PythonSQLRestMachine Learning
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

AWSAzureGCPRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About Role We are looking for an Operations Program Manager - Robotics Data Acquisition to own the day-to-day operating rhythm in our data collection facilities. You will work closely with operators, technicians, program managers, and engineers to keep rigs ready, campaigns moving, issues resolved, and performance improving. This is a hands-on operations role that requires you to be comfortable spending time on the floor, working through ambiguity, and using data to make the operation more reliable and efficient. This role is based in San Francisco, CA and requires in-person presence 5 days a week. In this role you will: Coordinate daily operations readiness across workstations, operators, materials. Track core operating metrics including utilization, cycle time, throughput, downtime, operator productivity, and data quality. Identify bottlenecks through workflow analysis, time studies, and capacity modeling, then drive practical fixes. Execute the rollout of new hardware, sensors, tools, and process changes with Engineering, Operations, Facilities, Supply Chain, and Safety. Identify equipment readiness issues and coordinate with technical support to keep workstations, and test equipment calibrated, configured, maintained, and ready for rollouts and evaluations. Lead root cause analysis for recurring operational issues and follow through on corrective actions. Provide operation input to create and maintain SOPs, work instructions, training materials, and process controls. Identify and flag resource constraints and manage issue escala

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

$226K – $285K/yr

Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a Supply Chain Program Manager, you will own material readiness and supply chain execution for critical hardware programs spanning custom silicon, systems, memory, storage, networking, and rack infrastructure. You will work cross-functionally with Engineering, Strategic Sourcing, Manufacturing Operations, Finance, Planning, Quality, and external suppliers to develop and execute scalable supply strategies that support aggressive product development and deployment timelines. This role requires deep understanding of hardware supply chains, material planning, NPI execution, supplier management, and operational scaling in constrained and rapidly evolving environments. In this role you will: Material Readiness & Supply Planning - Own end-to-end material readiness across NPI and production phases, including building the necessary framework and processes for enablement. Drive supply planning and execution for long lead-time and constrained commodities including ASICs, HBM, DDR, SSDs, networking, optics, power, thermal, and mechanicals. Build and manage material readiness plans aligned to proto/pre-EVT, EVT, DVT, PVT, and mass production schedules. Monitor supply health, lead times, inventory positions, allocation risk, and capacity constraints. Drive shortage management, allocation mitigation, and recovery planning. Coordinate supply commits, forecast alignment, and supply continuity planning with suppliers and manufacturing partners. Cross-Functional Program Ma

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C

SQLAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Support team is central to ensuring that our customers' experience with our products is nothing short of exceptional. We resolve complex issues, provide technical guidance, and support customers in maximizing value and adoption from deploying our products. We work closely with Sales, Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. OpenAI's customers represent a range of diverse backgrounds and maturity, from individual customers to early-stage startups and established global enterprises. Given OpenAI’s breakneck shipping cadence and growth – and the expectation that it will only accelerate – our ability to architect automation systems and agentic workflows for scale is central to our ability to maintain exceptional support quality in the face of AGI. About the Role We are seeking a Support Operations Lead who combines operational leadership, systems thinking, vendor management, and hands-on execution. You’ll own service health, automation programs, partner and vendor management. In addition to delivering high-quality service, you’ll identify opportunities to reduce manual work, experiment with tools and help operationalize AI across support at scale. This is not a traditional support operations lead role. We’re looking for someone to help us define the future of support, who thrive at the intersection of team/project management, systems building, data science/engineering, and with deep craft experience in the support operations space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead and evolve organizational design for our frontline operations across partner and vendor management, coaching teams to expand automation and deliver measurable capacity gains. Lead multi-site partner management, including commercial ownership, capacity planning and workfor

AWSRestAIGo
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%
Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We’re looking for a Finance Business Partner to join Notion’s Strategic Finance team. In this role, you’ll partner closely with leaders across Finance, People, Legal, and other business functions to drive financial insights, improve planning processes, and provide strategic analyses to enable the company to scale thoughtfully. The ideal candidate brings deep FP&A expertise and a passion for automation and continuous improvement. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Partner with executive leaders and functional leaders to drive critical business decisions. Build trusted relationships and proactively surface insights, risks, and opportunities. Build and own financial models, including capacity and benefits modeling. Prepare strategic trade-off analyses that help leadership prioritize investments so Notion scales responsibly. Partner with Corpo

SQLRestAIGo
E(
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $250K/yr

Quick readStrong listing-quality and freshness signals

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Skills and Qualifications: Product Knowledge: Deep understanding of product offerings and their real-world applications, particularly in AI technologies, enabling effective communication of value to potential customers. Sales Expertise: Proven track record in building positive relationships, prospecting, and negotiating high-value deals with senior executives. Experience generating pipeline through cold calling and prospecting. Interpersonal Skills: Strong communication and interpersonal abilities, with the capacity to be personable yet persistent, adapting your approach to different customer needs. Technical Acumen: Effective combination of selling skills and technical understanding, focusing on strategic thinking, problem-solving, and value creation tailored to each unique client. Adaptability: Ability to thrive in a fast-paced, dynamic environment, with a high level of flexibility and the capacity to adapt to changing priorities and overcome challenges. Responsibilities: Lead Generation: Responsible for outbound prospecting, generating new leads, and developing comprehensive strategies to expand the company’s presence within targeted institutions or regions. Client Engagement: B

🔔

Get new capacity strategy and operations jobs in United States by email

Daily job updates · Unsubscribe anytime