About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
Jobs in United States
Lead Cloud Operations Engineer in United States
2,434 active opportunities · Updated October 2026
Showing
15 jobs
Explore current lead cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. ABOUT THE TEAM Supply is responsible for knowing everything happening in the compute market: who's building, who's buying, and on what terms. This role owns a specific and fast-moving slice of that map — emerging clouds and international markets — and owns the full relationship lifecycle in that space, from first outreach through to closed terms. RESPONSIBILITIES Build and maintain a real-time picture of the emerging cloud and international compute landscape — who's active, what they're building, and what terms are available Own the full partnership lifecycle in this space — from identifying and sourcing new providers, to negotiating terms, to ongoing relationship management Develop and manage relationships across a broad set of emerging and international providers, from account reps up through leadership Identify, structure, and help close opportunities where Baseten can move quickly to secure favorable capacity terms Define compelling value propositions tailored to different types of providers, rather than a one-size-fits-all pitch Partner closely with others in the team already covering this space to build out a durable, well-organized intelligence and relationship function Collaborate with the broader Supply and Deals functions to bring opportunities to the table and support negotiation when it's time to close WHAT WE’RE LOOKING FOR Equal parts relationship-builder and operator — you can open a door and also drive it
We are looking for a Senior Forward Deployed Engineer to join the Customer Solutions team. You will be the technical authority embedded with our most complex customers, guiding them through deployment, architecture, onboarding, and the adoption of agentic development workflows. You bring deep hands-on experience from prior roles and use that depth to advise, design repeatable patterns, and drive customer outcomes end to end. You operate autonomously, own the technical success of your customers, and bring their experience back to shape how Coder builds and delivers. This is not an execution-only role. You are equally comfortable doing deep technical work with a customer and stepping back to design the repeatable pattern behind it. You are energized by ambiguity, motivated by customer outcomes, and capable of influencing organizational change alongside the technical work. This position is required to sit in the Eastern Time Zone. What You'll Do Serve as the primary technical authority for post-sales customers, guiding deployment architecture, environment design, and adoption of Coder across both human and AI development workflows Own onboarding engagements end to end, ensuring customers move from contract to productive adoption with speed and confidence Lead Get Well engagements where architecture decisions, rollout patterns, or organizational dynamics are limiting customer health or growth Help customers implement the technical and organizational changes required to adopt agentic development practices at scale Design and document repeatable delivery patterns across onboarding, architecture, and adoption that can scale across customer segments Design and recommend reference architectures tailored to each customer's cloud environment, security posture, and organizational constraints Translate customer environment complexity into clear guidance on networking, ingress, identity, and infrastructure patterns Anticipate technical and operational risks, escalate to the right
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 We are seeking a highly technical and experienced Engineering Manager to one of our business critical engineering team. This role is ideal for a hands-on leader who thrives in a fast-paced environment, has a deep understanding of collaborative document editing technologies, and is passionate about building scalable, high-performance software. As the Manager, you will oversee the development and delivery of innovative features, ensure technical excellence, and mentor a team of talented engineers. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our product. Operational Excellence : Ensure the stability and performance of the Docs platform, proactively addressing technical debt and optimizing system architecture. Qualifications: Technical Expertise : Proficiency in
$185K – $220K/yr
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We are seeking a strategic and technically fluent Lead, IT Audit to join our Finance team reporting to the Head of Internal Audit. This is a broad, high-impact role spanning both IT SOX compliance and operational IT audits. You will help establish and elevate our technology controls program end to end — owning the IT SOX lifecycle, designing the IT general and application controls framework, embedding AI and automation into how we test and monitor controls, and delivering value-added operational IT and cybersecurity audits that strengthen how the company builds and runs its systems. You will partner with leaders across Engineering, Security, IT, Finance, and the business to ensure sound technology controls are built into how the company operates as we scale. This role is ideal for someone who thinks like a builder, not just an auditor — someone who can translate complex control and security requirements into practical, scalable processes in a fast-moving SaaS environment with modern cloud architecture and complex data flows. This role can be based in either San Francisco or New York City. We work from our offices on M
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. In this role you will: Develop interactive, data-rich user interfaces using React, TypeScript, and Vega, with a focus on integrating LLM-driven features (e.g., natural language querying, generative UI, and AI-assisted data storytelling). Lead the end-to-end delivery of substantial product features, ensuring AI outputs are presented with high reliability and low latency. Work closely with PMs, UX designers, and AI/ML engineers to bridge t
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Staff Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A STAFF SOFTWARE ENGINEER, YOU WILL: Design and develop features, understand customer requirements and meet business goals. Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. Build highly reliable and fault-tolerant software to meet the needs of the largest customers. Ensure operational readiness and maintainabilit
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p
From $186K/yr
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity Are you ready to step into a pivotal leadership role where your engineering depth directly shapes the future of our core platform? As our new Engineering Manager, you will lead a talented, distributed team across US and EU time zones, acting as the critical manager bridging regional collaboration. Our Cloud Foundation team is the backbone of the New Relic platform. In this role, you won't just manage tasks; you will mentor and empower engineers, transitioning our operational framework from a reactive state to a culture of proactive ownership and engineering excellence. You will oversee critical global initiatives, including major regional expansions into FedRAMP High / IL4, India, and Australia. If you thrive on solving complex multi-cloud challenges at an exabyte scale while helping engineers grow in their careers, this is your opportunity to make a lasting impact. What you'll do Empower & Mentor: Lead and nurture a high-performing engineering team across the US and EU, facilitating career development, performance growth, and a collaborative team culture. Drive Strategic Ownership: Champion a shift from reactive delivery to proactive technical ownership, establishing best practices for platform reliability and cross-regional alignment. Lead Regional Expansions: Architect and execute key global infrastructure expansions across complex environments (including FedRAMP High / IL4, India, and Australia). Architect for Extreme Scale: Guide decisions around micr
$170K – $250K/yr
At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. The Role You'll Play At Mindbody, Core Engineering builds and evolves the foundational systems that help our products run reliably at scale. In this staff role, you’ll bring clarity to complex technical problems, guide architecture, and strengthen how we design, deliver, and operate the backend services that power real-world experiences. Lead cross-team technical execution, aligning architecture and delivery across multiple squads and core domains Design and evolve microservices patterns that improve reliability, performance, and maintainability Drive cloud and deployment improvements across AWS, Mindbody’s cloud platform, and our containerized deployment environment Partner with engineering and product leaders to turn ambiguous problems into clear technical plans and milestones Establish and socialize standards for service design, APIs, and relational data modeling (SQL) Strengthen monitoring and operational visibility using New Relic and Kibana, turning insights into durable system improvements Mentor and unblock engineers through design reviews, pairing, and practical guidance Reduce technical risk and complexity while balancing
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Senior Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A SENIOR SOFTWARE ENGINEER, YOU WILL: ● Design and develop features, understand customer requirements and meet business goals. ● Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. ● Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. ● Build highly reliable and fault-tolerant software to meet the needs of the largest customers. ● Ensure operational readiness and ma
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a talented and passionate Senior Software Engineer for our Snowpark Container Service , part of our Snowflake Compute Platform - to build our elastic, high-scale, high-performance, cloud native compute platform to enable bringing Compute to Data effortless and simple. Snowpark Container Services is a fully managed container offering that helps our customers easily deploy, manage, and scale containerized applications without having to move data out of Snowflake. As a fully managed service, it comes with Snowflake security, configuration, and operational best practices built in. You will be part of this highly productive, fast moving, and growing team that is critical to realizing Snowflake’s Data Cloud Mission. AS A SENIOR SOFTWARE ENGINEER, YOU WILL: ● Design and develop features, understand customer requirements and meet business goals. ● Lead a team of engineers, including mentoring and guiding them, and build technical direction and strategy for large and critical parts of the product surface area. ● Manage all aspects of the Project, including Design, Coding, Reviews, Testing, Observability, Tooling and On-Call support. ● Build highly reliable and fault-tolerant software to meet the needs of the largest customers. ● Ensure operational readiness and ma
About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. Technical Program Managers at OpenAI play a key leadership role in scaling these efforts, partnering deeply with product, engineering, design, and go-to-market teams to bring ambitious ideas to life and ensure clarity and discipline in execution. About the Role We are hiring a Technical Program Manager to support OpenAI's critical AI deployments across strategic cloud partners. This role is designed for a candidate who can operate as an end-to-end owner across internal engineering teams and external partner organizations. This role will drive the technical strategy and execution required to bring OpenAI models and platform capabilities into partner environments responsibly and at scale. The work spans engineering deliverables, shared roadmaps, model launch pipelines, technical integration, launch readiness, and post-launch follow-through. You will work closely with senior leaders across OpenAI engineering, infrastructure, product, safety, security, legal, finance, and go-to-market, as well as technical counterparts at our partners. The job is to turn broad partnership commitments into concrete execution plans, align both sides on what must land, and build repeatable mechanisms for launching OpenAI capabilities on third-party platforms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end execution for major cloud partner programs spanning model deployment, product integration, operational readiness, launch follow-through, and partner-platform adoption. Own integrated technical roadma
About the Team OpenAI’s Stargate and 3P Engineering teams are responsible for building and scaling the external infrastructure ecosystem that powers advanced AI systems. We work across hyperscalers, colocation providers, cloud partners, and strategic third-party operators to turn contracted capacity into production-ready compute. Our scope spans the full lifecycle of external deployments: commercial alignment, technical readiness, network integration, hardware enablement, operational readiness, and long-range scaling strategy. As OpenAI’s infrastructure footprint expands globally, we need leaders who can convert complex partner environments into reliable, high-velocity capacity for training and inference workloads. About the Role We are seeking a Technical Program Manager, Token-as-a-Service (TaaS) to lead delivery of external compute capacity that directly serves OpenAI model workloads. In this role, you will own complex cross-functional programs that transform third-party infrastructure into usable tokens at scale. You will partner across engineering, capacity planning, networking, hardware, finance, product, and external providers to ensure that deployed capacity translates into real production throughput. This role sits at the intersection of infrastructure execution, systems readiness, and business impact. Success requires strong technical fluency, elite program management, and the ability to drive accountability across internal teams and external partners. This is a high-visibility role with direct impact on OpenAI’s ability to scale model training and inference globally. This role is based in San Francisco, CA, with a hybrid work model of 3 days in office per week. Relocation assistance is available. Key Responsibilities Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply. Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments. Build
Other cities to consider
More places hiring for this role
Get new lead cloud operations engineer jobs in United States by email
Daily job updates · Unsubscribe anytime