Jobiba hiring network

Cloud Operations Engineer Jobs

2,329 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
11 days ago

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Baxter is seeking an experienced DevOps Engineer to support enterprise cloud platforms that securely connect medical devices and clinical applications with Baxter and third-party systems. This role will design, automate, deploy, and support cloud infrastructure across multiple environments. The successful candidate will bring strong technical skills, personal ownership, and the ability to collaborate effectively within a regulated healthcare environment. Key Responsibilities: Design, deploy, and maintain Azure infrastructure using Terraform and infrastructure-as-code principles. Build and support Azure Kubernetes Service (AKS) infrastructure, including clusters, node pools, namespaces, workloads, resource configurations, ingress, networking, and scaling. Develop and maintain Helm charts, Kubernetes manifests, and environment-specific configurations. Develop and maintain secure CI/CD pipelines using Azure DevOps, GitHub Actions, and related automation tools. Support Azure services including PostgreSQL Flexible Server, Cosmos DB, Az

pythonpostgresqlredis
View job →
MI
Mitsogo Inc
📍 Atlanta• Full-time
15 days ago

About Hexnode Hexnode, the Enterprise software division of Mitsogo Inc., was founded with a mission to simplify the way people work. Operating in over 100 countries, Hexnode UEM empowers organizations in diverse sectors. Fueling the transformation to a seamless ecosystem of connected tools, Hexnode is revolutionizing the enterprise software and cybersecurity landscape. Role Overview We are seeking a AWS Operations Specialist to manage and maintain our cloud infrastructure and device ecosystems. This is a highly operational, execution-focused role—not an architecture position. The ideal candidate has 2 to 4 years of experience executing infrastructure as code, monitoring environments, and following documented playbooks to keep our systems secure and resilient. Because this role handles secure environments, candidates must be US Citizens and capable of passing a comprehensive federal background check. Key Responsibilities Infrastructure Execution: Run, maintain, and execute existing Terraform and Ansible scripts to deploy and update infrastructure. GovCloud Monitoring: Actively monitor our AWS GovCloud dashboards, keeping a close eye on system health, performance metrics, and security baselines. Mobile Device Management: Manage Android Enterprise kiosk configurations, ensuring secure deployments and smooth device operations. Incident Response & Triage: Respond swiftly to operational alerts by strictly following our documented team playbooks. Escalation: Identify anomalies or issues that fall outside established, documented procedures and escalate them accurately to the engineering team. Required Qualifications & Profile Citizenship: Must be a US Citizen (required for GovCloud infrastructure management). Background: Must be able to successfully clear a rigorous federal background investigation. Experience: 2 to 4 years of hands-on experience in a technical operations, DevOps, or SysAdmin role. Technical Familiarity: Comfort executing/running Terraform and

awsaiswift
View job →
XH
XO Health Inc.
📍 India• Full-time• Remote• From ₹20L/yr
15 days ago

XO Health believes healthcare is fixable. Become part of the community changing the face of the industry. XO Health is the first health plan designed by and for self-insured employers that delivers a more unified health experience for everyone – from those who receive care, to those who deliver it, to those who pay for it. We are growing a multi-disciplinary team of diverse and digitally empowered employees ready to rebuild trust in healthcare through comprehensive and unified transformation. CyberSecurity & Infrastructure Engineer - India (Remote) About the Role : The Cybersecurity & Infrastructure Engineer is responsible for designing, implementing, monitoring, and securing the organization's hybrid cloud and infrastructure environments. This role serves as a technical leader for cybersecurity operations, cloud security, compliance initiatives, infrastructure engineering, and incident response. The position combines hands-on infrastructure administration with cybersecurity engineering responsibilities across Microsoft Azure, Microsoft Sentinel, Microsoft 365, Entra ID, AWS, networking, endpoints, and security platforms. Engineering plays a key role in maintaining compliance with SOC 2 Type II controls, improving cyber resilience, supporting audits, and advancing the organization's security maturity. Responsibilities: SOC2 Audit Support organization's annual SOC 2 Type II audit program, including control design, evidence collection, remediation management, auditor coordination, and continuous compliance monitoring. Partner with business and technology stakeholders to ensure security, availability, confidentiality, and change management controls are effectively implemented and operating throughout the audit period. Drive successful completion of SOC 2 Type II examinations with minimal findings by maintaining an audit-ready environment, strengthening internal controls, and promoting a culture of security and compliance. Develop and maintain policies, pr

REMOTEawsazuregit
View job →
SL
15 days ago

Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite

pythonjavareact
View job →
SL
15 days ago

Title: Senior Site Reliability Engineer - I, Product Area Focus Location: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering teams. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedit

pythonjavareact
View job →
SL
15 days ago

Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite

pythonjavareact
View job →
SC
Sigma Computing
📍 San Francisco• Full-time• $170K – $235K/yr
15 days ago

About the Role Sigma Computing is redefining business intelligence by making complex data analysis accessible through a high-performance platform built for the modern data stack. The Compiler Team plays a foundational role in this mission by transforming user-driven spreadsheet interactions into highly optimized SQL queries, enabling seamless exploratory analytics on cloud data warehouses. As a member of the Compiler Team, you will join a group of engineers dedicated to building the core systems and abstractions that power Sigma’s intuitive spreadsheet interface, ensuring speed, reliability, and scalability for all users. What You Will Be Doing Tackle core challenges at the intersection of data modeling, query compilation, and large-scale interactive analytics—making it possible for end-users to query data warehouses efficiently without deep technical knowledge Design, build, and maintain sophisticated compiler infrastructure and intermediate representations that translate spreadsheet operations into optimized query plans Apply advanced optimization strategies to improve performance and accuracy across a wide range of query workloads and data architectures Contribute to both backend (Rust) and key frontend foundations (TypeScript), evolving critical abstractions that enable end-to-end workflow optimizations and new features Debug, analyze, and resolve complex issues, ensuring robustness and maintainability in a rapidly evolving product Collaborate with engineers and product stakeholders to review designs and code, driving technical best practices and architectural decisions throughout the team and company Qualifications We Need 5+ years experience engineering high-quality software systems Demonstrated success building and maintaining complex infrastructure or core platform services Deep understanding of Computer Science fundamentals, particularly in compilers, algorithms, SQL Optimization Passion for teamwork, technical ownership, and continually

typescriptpythonsql
View job →
A
Addepar
📍 Pune• Full-time
15 days ago

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are currently seeking a highly experienced Staff Software Engineer with a strong Java background to join our Letter of Authorisation (LOA) team! We’re building a system that can self-service onboarding of held-away accounts for enterprise clients. This new system will be built on our application framework with cloud-native architecture principles. We’re looking for an experienced, detail-oriented Engineering Leader who will build an inclusive team culture, empower engineers to succeed and foster an environment that creates high quality engineering processes and product delivery. You are passionate about technology and can amplify your technical knowledge via your team. You’ll work closely with our product and design teams to deliver great products for our clients. Our engineering team primarily works in Java, Python and React, a strong Java experience with any of the additional skills would help succeed in the role. We use Agile methodologies to deliver impactful business outcomes. What You’ll Do Architect, implement, and maintain engineering solutions to solve complex problems; write well-designed, testable code. Lead individual project priorities, deadlines, and solutions. Collaborate effectively with product manager

pythonjavareact
View job →
A
15 days ago

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are currently seeking a Staff Software Engineer, Infrastructure to join the AI Platform team that powers seamless insights and interaction through natural language and data intelligence across our AI products. As a Staff Software Engineer, you’ll architect, build, and operate the backend and platform systems that power AI Platform. You’ll work across service design, distributed systems, cloud infrastructure, event-driven processing, observability, CI/CD, and production reliability, helping shape the technical direction of a platform that supports scalable, client-facing AI experiences. This role requires a strong software engineering foundation combined with deep infrastructure and systems thinking. We are looking for an engineer who can write high-quality production code, make sound architectural tradeoffs, and own platform capabilities end-to-end — not someone focused only on scripting, cloud configuration, or infrastructure tooling in isolation. You will collaborate closely with frontend, product, and AI/ML engineers to deliver reliable, secure, and scalable systems that align with Addepar’s standards of performance, resilience, and trust. Applicants must have legal authorization to work in the country where this role is based o

pythonjavasql
View job →
RS
Redwood Software
📍 Hyderabad• Full-time
15 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Software Engineer, Platform & Integrations . You will own the end-to-end delivery of complex features, optimize backend performance, and collaborate on architectural decisions that impact more than 1,000 enterprise customers worldwide. This is an ideal role for an experienced engineer who thrives on technical autonomy, enjoys tackling complex cloud-native challenges, and wants to have a direct impact on product direction. Feature Ownership & Architecture: Design, build, and maintain scalable, secure, and highly observable backend services and microservices using Java and Spring Boot. Platform Evolution: Actively contribute to upgrading our core platform infrastructure, focusing on system resilience, performance tuning, and seamless component communication. Security & Compliance: Implement rigorous secure coding practices to safeguard data exchange and ensure compliance across our multi-tenant SaaS environment. Collaborative Execution: Work closely with Product, QA, and senior le

javasqlpostgresql
View job →
S
16 days ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Proactive Threat team is responsible for identifying vulnerabilities and security weaknesses across Stripe's systems, applications, networks, and cloud infrastructure — before adversaries do. We operate as a hybrid offensive function: conducting penetration testing, emulating real-world threat actors through red team operations, and partnering closely with our defensive security teams to validate detection capabilities and improve Stripe's overall security posture. We are builders first. Our team develops custom tooling, automation frameworks, and internal platforms that scale our offensive capabilities and enable repeatable, high-fidelity assessments. We believe the best offensive security engineers are equal parts hacker and engineer. The team is distributed across the United States, primarily operating in Eastern and Pacific time zones, and collaborates regularly with security, engineering, and product stakeholders across Stripe — including teams in Europe and Asia. What you'll do As an Offensive Security Engineer on the Proactive Threat team, you will simulate the tactics, techniques, and procedures (TTPs) of real-world adversaries to uncover security risks across Stripe's products and infrastructure. You'll conduct hands-on penetration testing, lead red team engagements, and collaborate with blue team counterparts to validate and improve detection and response capabilities. Your work will directly influence how Stripe builds, ships,

pythonawsazure
View job →
O
Okta
📍 Bengaluru• Full-time
18 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we are building the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a individual contributor role for a hands-on system engineer who executes moderately complex tasks, builds platform components and collaborates under senior guidance. What You'll do : Contribute technical design and execution for automation initiatives within the team, creating paved paths that enable builders across Okta to connect enterprise systems seamlessly. Design, build, and deploy high-throughput event-driven integration flows , API gateways, and async orchestration workflows using AWS architectures and iPaaS platforms. Build reusable frameworks , developer SDKs, self-service primitives, and integration templates to streamline automation delivery. Develop and maintain AWS cloud services and modern iPaaS tooling, driving scalable architecture, reliability, and builder enablement. Partner with operations, security, and platform teams to strengthen monitoring, observability, structured audit logging, and automated governance. Actively participate in code reviews, team agile ceremonies, and technical discussions Collaborate with security, compliance, and business stakeholders to ensure self-service automations are secure, resilient, zero-tr

typescriptpythonjava
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Principal Software Engineer - PAM Okta is the identity standard. The Okta Identity Cloud is an independent and neutral platform that securely connects the right people to the right technologies at the right time. We help organizations do two things - secure and manage their extended enterprise, and transform their customers' experiences. With over 14,000 customers, 7000+ app integrations, and well over 200 million registered users, we are only getting started. The Okta Privileged Access Management (PAM) is an identity-centric approach to a common and critical privileged access use case. Our elegant Zero Trust architecture is purpose-built for the modern cloud and helps customers solve challenging security and operations pain points at scale. We're looking for a staff-level fullstack engineer to join a team of highly skilled and talented team players who're proud of what they own and deliver. Our elite team is fast, creative, and flexible; with a weekly release cycle and individual ownership, we expect great things from our engineers and reward them with stimulating new projects, new technologies, and the chance to have significant equity in a company that is changing the cloud computing landscape forever. You Will Leverage cutting-edge AI pair-programmers and LLMs (such as Copilot and Claude) to accelerate the development of secure, enterprise-grade Privileged Access Management (PAM) products. Proven expertise leveraging AI devel

javascriptjavareact
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

kubernetesgitmachine learning
View job →
M
Modal
📍 New York• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track

aigorust
View job →
🔔

Get new cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More cloud operations engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.