Jobs in United States

Cloud Operations System Administrator in United States

698 active opportunities · Updated October 2026

Explore current cloud operations system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 Woonsocket, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: We are seeking a highly skilled backend-focused Senior Software Engineer to join our modernization team focused on transforming the legacy CVS pharmacy systems into a cutting-edge, cloud-native platform. As part of the team, you will play a crucial role in designing and developing microservices that drive the modernization of critical patient and drug domains. You will work on integrating cloud-native solutions, enhancing performance, and ensuring seamless operation within a highly scalable and secure environment. The ideal candidate will have extensive experience in backend development, system design, and a strong understanding of cloud-native software engineering principles. Responsibilities: · Design, build, and maintain scalable data pipelines to support analytics, ML, and operational reporting · Develop robust data ingestion, transformation, and integration workflows using Python, SQL, and modern data engineering frameworks · Build and maintain batch and streaming data pipelines leveraging technologies such as Kafka (or similar pub/sub tools) · Work with Google Cloud Platform (GCP) services, including Cloud Storage, Dataflow, Pub/Sub, BigQuery, Cloud Spanner and Cloud Functions · Develop and manage data APIs and interfaces (REST and GraphQL) to enable high-performance data access across microservices · Implement CI/CD automation fo

PythonJavaSQLAWS
B
📍 Raleigh, North Carolina, United States
✓ High-confidence listingCompany trend +350%
Quick readStrong listing-quality and freshness signals

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Baxter is seeking an experienced DevOps Engineer to support enterprise cloud platforms that securely connect medical devices and clinical applications with Baxter and third-party systems. This role will design, automate, deploy, and support cloud infrastructure across multiple environments. The successful candidate will bring strong technical skills, personal ownership, and the ability to collaborate effectively within a regulated healthcare environment. Key Responsibilities: Design, deploy, and maintain Azure infrastructure using Terraform and infrastructure-as-code principles. Build and support Azure Kubernetes Service (AKS) infrastructure, including clusters, node pools, namespaces, workloads, resource configurations, ingress, networking, and scaling. Develop and maintain Helm charts, Kubernetes manifests, and environment-specific configurations. Develop and maintain secure CI/CD pipelines using Azure DevOps, GitHub Actions, and related automation tools. Support Azure services including PostgreSQL Flexible Server, Cosmos DB, Az

PythonPostgreSQLRedisAzure
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

KubernetesGitMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc

AWSGitRestAI
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About our Team: Micron’s Industrial and Physical AI team is driving the transformation of semiconductor manufacturing through Autonomous Operations, AI, robotics, and digital twin technologies! We develop and deploy innovative solutions across Micron’s global fabrication and assembly/test facilities, enabling smarter, safer, and more efficient operations at scale. Position Overview: We are seeking a hands-on Full-Stack AI Engineer to design, build, and deploy production-grade AI applications that support Micron's Autonomous Operations initiatives. This role owns the end-to-end development lifecycle, from data pipelines and AI models to APIs, web applications, digital twin integrations, and cloud/edge deployments, delivering impactful solutions for engineers, operators, and business leaders worldwide. Responsibilities: Design, architect, and deliver end-to-end AI products, including data ingestion pipelines, feature engineering, model training/inference, APIs, user interfaces, and application monitoring. Build and maintain modern front-end applications using React, Angular, or Streamlit, supported by backend services in Python and FastAPI. Develop scalable integrations between manufacturing systems, robotics platforms, AMRs, sensor networks, and enterprise applications to enable intelligent factory operations. Design and implement digital twin environments using platforms such as NVIDIA Omniverse, Gazebo, or Unity Robotics Hub to support simulation, validation, and o

JavaScriptTypeScriptPythonReact
B
📍 Raleigh, North Carolina, United States
✓ High-confidence listingCompany trend +350%
Quick readStrong listing-quality and freshness signals

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your Role at Baxter Provides enterprise-level technical leadership for cloud shared services and related connected-care ecosystems. Collaborates with engineering, product, and business leaders to shape long-term architectural strategy, establish technical standards, and guide the evolution of secure, cloud-native services that support multiple products, regions, and business domains. Serves as a senior technical leader for distributed systems, GraphQL and API architecture, multi-region cloud strategy, service interoperability, scalability, security, resiliency, observability, SOC 2 readiness, and operational excellence. Partners closely with executive leadership, product management, cybersecurity, quality, regulatory, operations, and engineering teams to align technology investments, architectural decisions, and platform capabilities with business objectives and sustained growth. <span style="color:

Node.jsAzureKubernetesGraphql
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.2%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

AWSRestAIGo
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Cyber team works to make frontier AI safe, trusted, and transformative for developers and enterprises. This team is building the security foundation for Codex: the native controls that govern what Codex can access and do, and the interfaces that allow customers and security partners to inspect, constrain, approve, and respond to Codex activity. Our goal is to make Codex secure by default, governable by enterprises, and interoperable with the security products customers already trust . This extends the existing product direction around tenant-scoped tools, guarded actions, approval systems, and scalable partner interfaces. About the Role We are looking for a deeply technical Product Manager to help build Codex security controls and the partner ecosystem around them. This role focuses on securing Codex itself : how identity, permissions, tools, MCP servers, repositories, secrets, networks, and high-impact actions are governed across Codex products. You will also help define standard interfaces through which authorized customer and partner systems can provide security context, inspect activity, return policy decisions, receive telemetry, and initiate bounded responses. You will work closely with Codex product and engineering, OpenAI Security and Safety, enterprise customers, and partners across application security, identity, cloud security, data security, infrastructure, and security operations. In this Role you Will Build native security controls for Codex Partner with engineering, design, security, and safety teams to develop controls for: Identity, roles, permissions, and tenant isolation. Access to repositories, files, tools, MCP servers, secrets, networks, and infrastructure. Read, write, execute, and deployment authority. Human and policy-based approvals. Prompt-injection and untrusted-content defenses. Audit trails, provenance, stop conditions, revocation, and rollback. Help establish a graduated authority model in which local, read-only

AWSCI/CDRestAI
S
📍 United States Minor Outlying Islands, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Security Engineer for our Enterprise Security team. The threats facing modern organizations are evolving at an unprecedented pace — driven by AI, cloud complexity, and an ever-expanding attack surface. Our Enterprise Security team is responsible for protecting Snowflake, our employees, and our customers by designing, implementing, and maintaining robust endpoint security solutions at enterprise scale, defending not just endpoints and identities, but the AI-assisted workflows and intelligent systems that define how the modern enterprise operates. AS A STAFF SECURITY ENGINEER AT SNOWFLAKE, YOU WILL: Develop, maintain, and scale Snowflake's endpoint security solutions while actively contributing to architecture and strategy. This includes EDR, DLP, secure browser, MDM, and AI security solutions across a complex multi-cloud, multi-SaaS enterprise environment. Own the technical roadmap for endpoint security tooling as Snowflake rapidly grows. Build intelligent security platforms and automate security operations using AI and robust software engineering. Identify opportunities to reduce toil, increase detection fidelity, and accelerate response through thoughtful, scalable automation. Monitor security events, investigate incidents, and build real-time detecti

JavaScriptPythonJavaAWS
P
📍 San Francisco, CA, United States· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is seeking a Sr. Manager to lead our Capacity Engineering team. The team ensures that Pinterest’s cloud infrastructure has the capacity it needs while operating reliably, efficiently and with clear financial accountability. You’ll lead the full portfolio across forecasting and supply, capacity-management systems, compute and GPU efficiency, infrastructure data and governance and capacity operations. What you’ll do: Lead the Capacity Engineering team and establish its 12–18 month functional and technical strategy, roadmap and success measures tied to Infrastructure and company goals. Develop CPU and GPU forecasts and supply plans that account for workload demand, delivery constraints, cost and reliability requirements. Guide the design and delivery of capacity requests, reservations, entitlements, allocation policy and infra

KubernetesAIFinance
C-
📍 New York, NY, United States
✓ High-confidence listing

$220K – $290K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We're looking for Senior Fullstack Software Engineers to help build the next generation of CLEAR's mission and vision. Beyond verifying identity, we're creating a secure, networked digital identity that enables seamless experiences across travel, enterprise, healthcare, financial services, and beyond. As a Senior Software Engineer III, you'll own complex technical problems from design through deployment, partnering closely with Product, Design, Security, and Operations to deliver reliable, scalable solutions. We're looking for engineers with a strong builder mindset who thrive in ambiguity, take ownership, and enjoy turning ideas into production systems. Level and team matching (open roles across the three pillars that make up Technology at CLEAR: Core Identity, CLEAR1 , and CLEAR Travel ) will occur towards the end of our interview process. Tech stack overview: A brief highlight of our tech stack: Java / Kafka / Postgres AWS cloud What you'll do: Advance our capabilities across a wide array of industries and domains and gain hands-on experience with privacy, security, data modeling and a

JavaPostgreSQLAWSDocker
F
📍 Mclean, Virginia, United States
✓ High-confidence listingCompany trend -26.7%

From $32/hr

Quick readStrong listing-quality and freshness signals

At Freddie Mac, our mission of Making Home Possible is what motivates us, and it’s at the core of everything we do. Since our charter in 1970, we have made home possible for more than 90 million families across the country. Join an organization where your work contributes to a greater purpose. Position Overview: Want to create technology that matters while building your future and your professional network? If you’re getting your degree in Computer Science or a related technology area and looking to get hands-on, skill-building experience with a leader in the financial technology industry, Freddie Mac’s Technology Summer Internship program is for you! Through this internship, you’ll build on the knowledge and skills you’ve acquired in your undergraduate education – while getting real-world technology and innovation experience and making lifelong connections. Upon successful completion of the internship and your undergrad degree, you may be extended an opportunity to join our Technology Analyst Program – a 12-month cohort experience that springboards you into a fulltime technology career with the industry leader in the housing/finance industry! Our Impact Enterprise Operations &#43; Technology (EO&#43;T) helps run and transform Freddie Mac’s business — from building applications and modernizing cloud and data platforms to protecting systems from cyberthreats and keeping critical operations resilient. As an intern, you can contribute to work that improves how we serve customers and help make home possible for millions of families. We are accepting applications for this position until 10/16/2026 Your Impact As an intern, you will support exciting projects, including: Software Development/Application Support Participate on project teams to develop innovative, high-quality software solutions in an Agile development environment Help d

PythonJavaAIPower Bi
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and

AWSGCPCI/CDMachine Learning
G
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -97.9%

From $86.4K/yr

Quick readStrong listing-quality and freshness signals

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Senior Internal Auditor reporting to the Senior Manager, Technology Internal Audit, you’ll help GitLab assess risk and strengthen controls across a technology landscape that includes multi-cloud infrastructure, artificial intelligence and machine learning systems, and modern development practices. This USA-based role supports our Sarbanes-Oxley Act (SOX) program while partnering with Engineering, IT Operations, Security, and business teams to build controls that work in practice, not just on paper. You’ll execute technology audits, turn findings into practical improvements, and use data analytics, au

GitRestAgileMachine Learning
S
📍 Bellevue, WA, United States· Full-time
✓ High-confidence listingCompany trend -91.7%
Quick readStrong listing-quality and freshness signals

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Corporate Systems Engineering builds and operates the software platforms, integrations, and automations that power Smartsheet’s core business functions across Finance, Sales/GTM, and People & Culture. Our team owns mission-critical systems and workflows that enable how the company hires, sells, bills, pays, reports, and scales. We operate at the intersection of software engineering, enterprise platforms, and business-critical data, treating internal systems with the same rigor, reliability, and product mindset as customer-facing software. The Automation team builds human-to-system and system-to-system automations that reduce manual effort and friction across the business. We combine cloud-native services, agentic AI, and workflow orchestration to enable employees to interact with enterprise systems through intelligent, secure, and auditable automation. As a Senior Software Engineer I (Automation), you will lead the design, build, and operation of systems and workflows that directly support business execution at scale. You will own complex technical initiatives, partner with Product Managers and stakeholders on technical roadmaps, and mentor junior engineers. This full-time position reports to the Sr. Director, Development and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Architect AI Agents: Take a leading role in designing Agentic Workflows using AWS Step Functions and Bedrock Agents that reason

JavaScriptTypeScriptPythonJava
🔔

Get new cloud operations system administrator jobs in United States by email

Daily job updates · Unsubscribe anytime