Jobs in United States

Cloud Operations Engineer in United States

698 active opportunities · Updated October 2026

Explore current cloud operations engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Job Summary We are looking for a GTM DevOps Engineer to join our Business Systems team and own the reliability, automation, and delivery infrastructure behind our Go-To-Market (GTM) technology stack. This role sits at the intersection of platform reliability and CI/CD engineering, ensuring that our critical business systems — including Salesforce, NetSuite, MuleSoft, Workato, and an expanding portfolio of AI-powered workloads — are deployed consistently, operate resiliently, and scale with the business. You will partner closely with Business Systems developers, architects, and business stakeholders to build and maintain the pipelines, monitoring frameworks, and operational standards that keep our GTM systems healthy and our release cycles fast and predictable. As our team builds and deploys AI agents across GCP Cloud Run and AWS Bedrock AgentCore, you will serve as the infrastructure and deployment owner for these workloads — bringing engineering discipline to an environment where AI-generated code is increasingly entering production. This is a hands-on engineering role for someone who thrives in complexity, takes ownership of platform uptime, and brings a software engineering mindset to business application operations — directly supporting GTMSOE's broader mission of operational excellence across the GTM org. Key Responsibilities CI/CD & Release Engineering Design, build, and maintain CI/CD pipelines for Salesforce (SFDX/Salesforce CLI), NetSuite (SuiteScript/SuiteBundler), MuleSoft (Anypoint Platform), and Workato; establish branching strategies, environment promotion standards, and release gatin

PythonNode.jsAWSGCP
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer leading Fleet Management, you will be the overall technical lead across three pods and the person who sets the technical direction for the fleet management layer of Roblox. This is a hands-on, deeply technical leadership role that owns all of Roblox's compute capacity end to end: from low-level provisioning and the data plane, up through the control planes that operate it, and all the way to the UI and internal-facing products that let teams self-serve capacity. Your org centralizes security, maintenance operations, and the uptime of every Roblox Kubernetes cluster, and governs the internal customer contracts that drive automation across the fleet spanning Roblox data centers and cloud providers. You will guide architecture, raise the engineering bar, and make sure compute capacity supply and demand stay in balance as the fleet grows. You will: Serve as the overall technical lead for three Fleet Management pods, setting and aligning the technical direction across low-level provisioning, the data plane, and the control plane and product surfaces above them. Architect the declarative, Kubernetes-style control planes that operate Roblox's compute fleet across o

SQLAWSKubernetesGit
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%

From $158.8K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Are you passionate about building and scaling enterprise finance systems that power global business operations? Join Pinterest’s IT Enterprise Systems team, where you’ll play a key role in evolving our Oracle EBS and Finance technology landscape. This is an opportunity to drive meaningful impact by partnering with Revenue, Finance,, and cross-functional teams to deliver reliable, scalable, and business-critical ERP solutions that support Pinterest’s continued growth. What You’ll Do: Design end-to-end solutions across billing, receivables, collections, subledger, and related finance processes to improve business resilience and operational efficiency. Architect and enhance integrations across Oracle EBS Financials, SaaS applications, and cloud platforms that support revenue operations, reporting, and downstream financial processes. Use AI thoughtf

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. About the Team The Information Security organization at Postman operates across three pillars: Governance Risk & Compliance (GRC), Product Security, and Security Operations. We are a team of builders, not checkbox-checkers. We hold active SOC 2 Type II, ISO 27001, ISO 42001, and HIPAA compliance postures, and we are pursuing FedRAMP High and CMMC Level 2 authorization. Our security stack includes Wiz, SentinelOne, Okta, Jamf, and 1Password, and we operate across a multi-cloud environment. The Offensive Security team is the "red" pulse of this organization. We don't just find bugs — we simulate the adversary to ensure our defenses hold up under real-world pressure. We focus on continuous security validation, AI-augmented adversary emulation, and offensive AI security research at Postman's scale. The Opportunity We are looking for a Principal Offensive Security Engineer who is as much a strategist as they are a hacker. You will own the strategic direction of Postman's offensive security program — including building out a dedicated Offensive AI Security capability from the ground up — and operat

AWSKubernetesCI/CDGraphql
P
📍 San Francisco, CA, United States· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Sr. Manager to lead our Capacity Engineering team. The team ensures that Pinterest’s cloud infrastructure has the capacity it needs while operating reliably, efficiently and with clear financial accountability. You’ll lead the full portfolio across forecasting and supply, capacity-management systems, compute and GPU efficiency, infrastructure data and governance and capacity operations. What you’ll do: Lead the Capacity Engineering team and establish its 12–18 month functional and technical strategy, roadmap and success measures tied to Infrastructure and company goals. Develop CPU and GPU forecasts and supply plans that account for workload demand, delivery constraints, cost and reliability requirements. Guide the design and delivery of capacity requests, reservations, entitlements, allocation policy and infra

KubernetesAIFinance
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Sr. DevOps Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. Scripting and

JavaScriptPythonJavaAWS
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Sr. Software Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary As a Software Engineer on this program, you will have the opportunity to build strong systems, software, and cloud environments while providing operations and maintenance for critical systems. This role will provide technical expertise in the design, development, implementation and testing of customer tools and applications. Based in a DevOps framework, this role participates in and/or directs major deliverables of projects through all aspects of the software development lifecycle including scope and work estimation, architecture and design, coding and unit testing. Primary Responsibilities: Participates in and/or directs software programming initiatives using Java, JavaScript, Python, SpringBoot, and Hibernate. Develops software system validation and testing methods using Junit and Katalon and uses integrated custom developed software solutions to leverage automated deployment technologies Develop, prototype and deploy solutions within Commercial Cloud Solutions leveraging infrastructure platform services Coordinate closely with team members, Product Owners and Scrum Masters to ensure User Story alignment and implementation to customer use cases Support th

JavaScriptPythonJavaAWS
B
📍 Raleigh, North Carolina, United States
✓ High-confidence listingCompany trend +350%
Quick readStrong listing-quality and freshness signals

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your Role at Baxter Provides enterprise-level technical leadership for cloud shared services and related connected-care ecosystems. Collaborates with engineering, product, and business leaders to shape long-term architectural strategy, establish technical standards, and guide the evolution of secure, cloud-native services that support multiple products, regions, and business domains. Serves as a senior technical leader for distributed systems, GraphQL and API architecture, multi-region cloud strategy, service interoperability, scalability, security, resiliency, observability, SOC 2 readiness, and operational excellence. Partners closely with executive leadership, product management, cybersecurity, quality, regulatory, operations, and engineering teams to align technology investments, architectural decisions, and platform capabilities with business objectives and sustained growth. <span style="color:

Node.jsAzureKubernetesGraphql
C-
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$145K – $170K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. CLEAR is seeking a Senior Security Operations Analyst III to join our SOC team to help strengthen our ability to detect, investigate, and respond to evolving security threats. In this role, you’ll lead complex investigations, improve CLEAR’s threat detection and response capabilities, and serve as a trusted security partner while helping develop the analysts and program around you. What you'll do: Lead complex investigations of security events across corporate networks, endpoints, data centers, cloud environments, and other critical systems, driving incidents from initial analysis through escalation and remediation Develop, tune, and optimize threat detection logic across SIEM, EDR, and other security platforms, proactively identifying coverage gaps, reducing false positives, and improving the fidelity of security alerts Partner with Engineering, Infrastructure, and other teams to investigate threats, identify root causes, communicate risk, and drive timely remediation and improvements to CLEAR’s security posture Apply threat intelligence, data, automation, and AI-enabled tools to identify emerging attack patterns, accelerate investigations, improve detection workflows, and strengthen decision-making while applying sound security judgment Serve as a subject matter expert and escalation point for other analysts, mentoring junior team members, sharing knowledge, and helping establish scalable processes, playbooks, and standards for threat detection and analysis Continuously evaluate CLEAR’s detection coverage against the evolving t

GitRestAIGo
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$275K – $350K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng

PythonJavaAWSKubernetes
H
📍 Louisville, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Humana is seeking a self-driven and collaborative Lead Engineer to join our Interactive Voice Response (IVR) team. In this role, you will deliver innovative IVR solutions and develop robust omnichannel APIs for our enterprise platforms. You will have the opportunity to drive the success of a high-impact, customer-facing application within a Fortune 50 company, working closely with multiple teams throughout the software development lifecycle (SDLC). Lead Engineer –Omnichannel Humana is seeking a self-driven and collaborative Lead Engineer to join our Omnichannel team. In this role, you will design, develop, secure, and enhance enterprise APIs that support high-impact, member facing, applications across Humana's digital and voice channels. This role offers the opportunity to modernize and strengthen existing API capabilities while helping deliver resilient, scalable, and secure omnichannel solutions within a Fortune 50 organization. Key Responsibilities Design, develop, and maintain scalable Omnichannel APIs that support enterprise applications and customer-facing capabilities. Enhance the security, resiliency, performance, and reliability of existing APIs through modernization, improved architecture, observability, testing, and operational controls. Apply AI and AI-assisted engineering practices to accelerate development, improve quality, automate testing, enhance documentation, and identify opportunities for optimization. Partner with architecture, security, cloud, product, engineering, and operations teams to deliver secure, resilient, and enterprise-aligned API solutions. Collaborate with agile teams to plan, track, and deliver API enhancements, platform improvements, and cloud-based capabilities. Develop proofs of

Machine LearningAIRecruitment
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc

AWSGitRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer you will join our OpenAI engineers and researchers in building, operating and securing transformational AI technologies. This role will focus on all aspects of Detection & Response but with a strong emphasis on detecting insider threats and influencing controls to safeguard OpenAI's most sensitive assets. In this role, you will: In this role, you will: Innovate on Detection and Response infrastructure to engineer and automate end-to-end detection and investigation workflows. Develop, measure, and tune detection rules to ensure effective and sustainable operations. Drive projects across OpenAI’s technology stack with a focus on insider threats, ranging from access abuse and intellectual property theft to novel risks emerging within AI infrastructure. Partner closely with cross-functional stakeholders, including HR, Legal, and peer investigative teams, providing technical expertise and evidence to support investigations. Collaborate on cutting-edge AI research, and use AI to improve OpenAI’s Security posture. You might thrive in this role if you: 5+ years experience working in a detection/response or insider-risk role.. We are seeking mid-level and senior candidates. You have broad familiarity with operating systems and platforms such as macOS, Windows, Linux, and Kubernetes, along with experience in cloud infrastructure. Knowledge of modern adversary tactics and attack paths, data exfiltration techniques, and h

PythonAWSKubernetesLinux
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product workloads. The company seeks a Senior Technical Program Manager (TPM) to lead complex, cross-functional programs powering NVIDIA’s next-generation AI software platforms. In this role, you will drive software initiatives across platform services, cloud infrastructure, and system integration. The focus is on enabling scalable, reliable, and supportable software for AI workloads. You will be responsible for managing high-impact engineering programs within a dynamic, fast-paced roadmap, aligning priorities across teams, and ensuring timely, high-quality delivery. This role requires strong technical competence, a proactive approach, and the ability to operate effectively across multiple levels of the organization. This is a software-first TPM role. The ideal candidate has extensive experience managing software initiatives. They also understand the full-stack environment, including infrastructure dependencies, system bring-up, integration readiness, and operational needs to support software across stack layers. What You'll Be Doing: Lead end-to-end execution of software platform initiatives, including planning, execution, delivery, and operationalization. Work together with software, infrastructure, product, and operations teams to ensure alignment on goals, deliverables, achievements, and schedules. Lead cross-functional initiatives encompassing cloud-native services, platform software, system integration, and release delivery. Help connect software roadmap execution to full-stack readiness, including dependencies across infrastructure, bring-up, validation, and downstream operational support. Identify cross-functional dependencies, mitigate risks, and drive resolution of complex technical and programmatic issues. Establish clear success metrics and reporting mechanis

KubernetesLinuxMachine LearningArtificial Intelligence
🔔

Get new cloud operations engineer jobs in United States by email

Daily job updates · Unsubscribe anytime