About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
Jobs in United States
Cloud Operations Lead in United States
698 active opportunities · Updated October 2026
Showing
15 jobs
Explore current cloud operations lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role A strong and reliable platform is essential to scaling Sentry for the future. Our Platform organization is responsible for everything that powers Sentry—from cloud infrastructure and streaming systems to storage, deployment, and security. We own the core services and technical foundations that enable every product and engineering team at Sentry to move fast and build with confidence. We're looking for a passionate and pragmatic Senior Staff Software Engineer to help lead this evolution. In this role, you’ll report directly to the VP of Engineering and collaborate with teams across the company to shape the future of Sentry’s platform. What You’ll Do Architect the future of Sentry by translating business needs and product strategy into clear, scalable technical blueprints. Partner with product and engineering leaders to align technical roadmaps with company goals. Lead cross-cutting initiatives across the Platform org—owning them end-to-end and driving meaningful outcomes. Promote engineering excellence by mentoring platform engineers, sharing best practices, and setting high standards for system design, scalability, and operational quality. Review major architectural proposals and help ensure consistency, maintainability, and long-term technical health across the company. You’ll Love This Job If You... Enjoy designing and building platforms that help teams move faster and scale safely. Thrive on solving complex, multi-dimensional problems across product, infrastructure, and organizational layers. Want to make architectural decisions that shape Sentry’s long-term success. Bring new ideas, tools, and frameworks t
About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage. Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved. Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more. We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves. We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role You will lead the Financial Services and Insurance industry strategy and business growth agenda, defining how the organization serves banks, credit unions, fintechs, payments companies, wealth firms, and insurers through industry-specific customer experience automation, cloud modernization, AI, and workflow orchestrati
From $124.5K/yr
About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage. Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved. Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more. We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves. We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role The ISV & Strategic Alliances Sr. Partner Marketing Manager leads high-impact partner marketing programs that drive measurable pipeline and revenue growth through Dialpad's strategic technology and alliance partnerships. This role primarily supports the Google Cloud partnership globally and T-Mobile for Business na
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: As a New Grad Software Engineer, you'll join a team of exceptional builders working on products that are reshaping how the world creates software. You'll have the opportunity to work on everything from our AI-powered development platform to the distributed systems that enable real-time collaboration for millions of developers. This is a chance to define your career while defining the future of software development. You'll work on problems that matter, with the autonomy to drive solutions and the support to grow into a technical leader. What you will build: Product features that delight users and make it possible for anybody to create software AI coding agent that understands intent and generates production-ready applications Cloud infrastructure that provides instant, powerful development environments at global scale Platform features that enable one click deployments and scale to millions of users Required skills and experience: Recent graduate (2027) with a degree in Computer Science, Computer Engineering, or related field Strong programming skills in a modern language (JavaScript/TypeScript, Python, Go, Rust) Full-stack capabilities with experience in React, Node.js, and database technologies Growth orientation - eager to learn new technologies and take on increasing responsibility Collaborative spirit - you work well in cross-functional teams and value diverse perspectives What we value : Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions Self-directed and autonomous: Capable of working independently while collaborating effectively with cross-functional teams Strong communication skills: Ability to explain complex technical conce
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a Delivery Director, Capacity programs for our on-premises data center builds and neo cloud (GPU cloud) delivery programs. This is a high-visibility, execution-critical role sitting at the intersection of infrastructure engineering, capacity planning, vendor/partner management, and customer delivery. You will own the end-to-end delivery lifecycle for large-scale compute infrastructure — from initial site/capacity commitments through power, networking, and hardware bring-up, to production-ready GPU/compute capacity landing in the hands of internal teams or customers. You'll be the person who turns ambitious infrastructure roadmaps into predictable, on-time, delivery. RESPONSIBILITIES Own delivery of on-prem infrastructure builds — colocation expansions, power/cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing — coordinating across colo providers and partners, network engineering, hardware ops, and vendor teams. Drive neo cloud delivery programs — manage capacity delivery from GPU cloud and neo cloud partners (e.g., colocation/bare-metal/GPU cloud providers), including contract milestones, capacity ramps, SLAs, and go-live readiness. Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power/shell timelines, hardware lead times, logistics, and software/platform readiness into a single critical path.
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team focused on transforming how Mastercard's payment systems are built, scaled, and operated. As a Senior Software Engineer, you will lead the design and development of cloud-ready applications, microservices, and distributed systems that support large-scale payment processing platforms while helping advance modernization, automation, and engineering excellence across the organization. In this role, you will contribute to software architecture decisions, drive technical design discussions, and partner with engineers to deliver scalable, resilient, and maintainable software solutions. You'll have the opportunity to solve complex technical challenges, mentor other engineers, and influence how software is designed, developed, tested, and supported across critical technology platforms. What You Will Do •Design software solutions and contribute to software architecture decisions that support scalability, maintainability, and operational excellence. •Translate complex product requirements into technical designs and implementation plans. •Lead development of modular, extensible, high-performance applications. •Design and implement comprehensive unit, functional, and integration testing strategies. •Analyze, optimize, and improve application performance, scal
From $186K/yr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity Are you ready to step into a pivotal leadership role where your engineering depth directly shapes the future of our core platform? As our new Engineering Manager, you will lead a talented, distributed team across US and EU time zones, acting as the critical manager bridging regional collaboration. Our Cloud Foundation team is the backbone of the New Relic platform. In this role, you won't just manage tasks; you will mentor and empower engineers, transitioning our operational framework from a reactive state to a culture of proactive ownership and engineering excellence. You will oversee critical global initiatives, including major regional expansions into FedRAMP High / IL4, India, and Australia. If you thrive on solving complex multi-cloud challenges at an exabyte scale while helping engineers grow in their careers, this is your opportunity to make a lasting impact. What you'll do Empower & Mentor: Lead and nurture a high-performing engineering team across the US and EU, facilitating career development, performance growth, and a collaborative team culture. Drive Strategic Ownership: Champion a shift from reactive delivery to proactive technical ownership, establishing best practices for platform reliability and cross-regional alignment. Lead Regional Expansions: Architect and execute key global infrastructure expansions across complex environments (including FedRAMP High / IL4, India, and Australia). Architect for Extreme Scale: Guide decisions around micr
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a VP of Finance to build the finance function from the ground up as our first full-time finance hire. This is a high-impact role for someone who thrives at the intersection of strategic thinking and hands-on execution. We are looking for someone who can architect the systems and processes that will scale with Modal, partner closely with the founders and executive team, and grow into the company's CFO. You'll report directly to the CEO and collaborate closely with our BizOps, GTM, and Product teams. In this role, you will: Build and maintain Modal's operating model, tying financial performance to company KPIs and resource allocation Lead all budgeting, forecasting, and long-range planning processes, and develop the reporting infrastructure that gives leadership and the board clear, timely visibility into the health of the business Partner with the found
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Requirements: 5+ years of experience writing high-quality production code Experience building high-performance distributed systems at a large scale (the more battle scars, the better) Strong cloud skills Strong knowledge of low-level operating system foundations (Linux kernel, file systems, containers, etc.) Experience with performance engineering (tell us a story of when you shaved off a few milliseconds!) Ability to work in-person in our NYC or SF office. Prior experience with Rust is nice to have, but not required. Ability to participate in on-call rotation and respond to production incidents.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At Modal, we sell cloud services atop which our customers run their critical production systems. As a rapidly growing new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform, customer base, and our team. This role is for people who are deep systems thinkers, love stacking nines, and thrive from making others move faster at scale. Responsibilities include: Identifying architectural changes to improve reliability and performance. Fostering a culture of reliability across Modal’s engineering organization. Defining and implementing operational processes such as deployments, upgrades, etc. Operating systems like Kubernetes, Postgres, Redis, etc. Participating in on-call rotations, and responding to production incidents. Requirements: 5+ years of experience writing high-quality production code. 2+ years of
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience in making ML systems performant at scale. If you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you! Requirements: 5+ years of experience writing high-quality, high-performance code. Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT). Familiarity with Nvidia GPU architecture and CUDA. Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc). Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
Other cities to consider
More places hiring for this role
Get new cloud operations lead jobs in United States by email
Daily job updates · Unsubscribe anytime