Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! Sr Solutions Engineer We are looking for a Sr Solution Engineer in Mumbai to support the rapidly growing Couchbase user community and help drive customer success. Our Solution (Pre-Sales) Engineers are the primary technical field experts, responsible for actively driving and managing the technical components of a sales engagement. This includes explaining Couchbase advantages and differentiators, and how our Couchbase solutions stack meets these customer challenges, all while getting customers excited about using this new approach for delivering consistent business value. This specific position focuses on assisting customers who interact with our technology via application code and the application development process. This development centric position requires experience in presales, software development e
Jobiba hiring network
Cloud Operations Engineer Jobs
2,329 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! Our Solution (Pre-Sales) Engineers are the primary technical field experts, responsible for actively driving and managing the technical components of a sales engagement. This includes explaining Couchbase advantages and differentiators, and how our Couchbase solutions stack meets these customer challenges, all while getting customers excited about using this new approach for delivering consistent business value. This specific position focuses on assisting customers who interact with our technology via application code and the application development process. This development centric position requires experience in Presales, software development experience as well as time developing business solutions. Experience with one or more Enterprise Programming Languages and Solution Development methodologies. Location: S
Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! Lead Software Engineer - Storage As a key contributing member of the storage development team, you will be responsible for enhancing the highly scalable and performant storage engines used by both Couchbase Server and Couchbase Capella. You will be developing a highly-available and concurrent enterprise-grade system software. Most of all, you will be able to celebrate the wins by experiencing the direct result of your hard work from our customers’ success stories. The ideal candidate will have a strong technical background, excellent communication skills, and proactive problem-solving skills. The innovative work storage team does has been widely recognized by the industry. The following publications in VLDB conferences reflect the storage work at Couchbase Nitro: A Fast, Scalable In-Memory Storage Engine for NoS
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team.... Global Compute runs Optimised Hosting, GoDaddy's global platform for all customer hosting products. Squad R is the engineering team responsible for operating, scaling, and continuously improving the OpenStack-based clouds that power that platform. We treat reliability as an engineering problem: we automate toil away, we plan capacity ahead of demand, and we instrument everything so that we understand our systems before they surprise us. As an SRE III on the team, you'll be a senior technical contributor who others lean on for the hard problems. What you'll get to do... Operate and scale GoDaddy's cloud infrastructure, including our OpenStack-based hosting platform. You'll troubleshoot and improve services spanning compute, networking, and storage in large-scale production environments. Drive the OpenStack migration. Help move customer hosting workloads onto the platform safely — designing and executing migration tooling, validation, and rollback strategies that protect customer experience. Work within a large-scale global hosting environment supporting thousands of servers and customer workloads across multiple regions. Eliminate toil through automation. Build and maintain automation in Python and Puppet to replace manual operational work. Treat repeated manual effort as a bug to be fixed. Strengthen observability. Improve monitoring, alerting, and dashboards so that signal reaches the right engineer at the right time, and so that we can reason about system behavior from data. Participate in on-call and incident response. T
Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. Senior Staff Security Engineer The Senior Staff Security Engineer owns current and target-state data architectures and reporting while also designing, implementing, and monitoring cloud (AWS/GCP) infrastructure security controls; deploying, securing, configuring, and operating SIEM and other security resources; identifying, triaging, and remediating infrastructure and web vulnerabilities; leading incident triage and external-researcher engagement; mentoring junior staff; and helping shape secure, scalable approaches for AI-enabled tooling, automation, and emerging product capabilities. Role summary and core responsibilities 6+ years of relevant experience Candidates must demonstrate proficiency in cloud security, network security, URL filtering, common security frameworks, and CVE lifecycle management Prac
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model. This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin. Responsibilities Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems Participate in a 24/7 on-call rotation to resolve critical issues Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice You may be a good fit if you Have 6+ years of experience in software development and operating distributed systems Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests) Have deep experience using and extending containerization technologies, preferably Kubernetes Have a solid understanding
About the Team The Codex team is responsible for building state-of-the-art AI systems that can write code, reason about software, and act as intelligent agents for developers and non-developers alike. Our mission is to push the frontier of code generation and agentic reasoning, and deploy these capabilities in real-world products such as ChatGPT and the API, as well as in next-generation tools specifically designed for agentic coding. We operate across research, engineering, product, and infrastructure—owning the full lifecycle of experimentation, deployment, and iteration on novel coding capabilities. About the Role As a Performance & Systems Engineer on the Codex team, you will be responsible for whole-system optimization across a complex, evolving stack. Codex spans LLM inference, cloud orchestration, agentic work management, and multiple product surfaces. Your job will be to identify and land high-leverage changes—across infrastructure, modeling, and product layers—that make Codex agents significantly faster and cheaper to serve. We’re looking for generalists who thrive in ambiguity and love chasing performance bottlenecks to ground. This is a high-ownership role where your work will directly improve the experience of millions of users. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond. Build tooling to measure, profile, and optimize system performance at scale. Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost. You might thrive in this role if you: Have experience operating across both ML systems and cloud infrastructure. Enjoy diving into messy, ambiguous problems and emerging with clear wins. Think holistically about performance, balancing spee
About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun
The Storage Layer Services team is currently re-architecting the MongoDB Cloud Storage Layer. This is a relatively new team in MongoDB that sits at the heart of the next generation MongoDB Cloud Storage Architecture, and the team is working to build performant multi-tenant distributed storage services both to enhance our existing MongoDB cloud storage architecture and to power more of our customers' use cases more efficiently. Engineering at MongoDB is globally distributed, with a mix of folks being fully remote, hybrid, or in-office. We have a small but growing team that calls Sydney home, and we are looking for a Staff Engineer to join the team working closely with other teams in Sydney and North America. Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies great engineering fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. We are looking to speak to candidates who are based in Sydney for our hybrid working model. You’re an ideal candidate if: You have 10+ years of experience in programming, debugging, and performance tuning highly concurrent and/or distributed systems. Especially if you have worked in a systems language (C, C++, Rust, etc) for a number of those years You have a track record as an effective technical leader. You love helping teams be successful at solving vaguely defined problems in iterative and measurable ways. You put the customer first, and don’t hesitate to cross team boundaries in search of the right solution You have a solid grasp of related systems fundamentals, such as cache management, log-based recovery, transactions or performance profiling You’re comfortable reasoning about highly concurrent, asynchronous services — backpressure, tail latency, and the failure modes of replicated state machines You’ve worked on large, highly availabl
MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. We are looking for a talented Senior Software Engineer to join our team as we execute on a multi-year roadmap and prepare to launch and rapidly scale services that handle petabytes of data. Come do some of the most interesting work of your career as we Think Big and Go Far for our customers! This role can be based out of our New York City office (hybrid working model) or remotely in the North America region. What you’ll do Design, build, and operate control plane services powering an elastic and multi-tenant storage layer for thousands of database instances. Solve problems around maintaining high availability and performance during load spikes, hardware failures, cloud provider outages, and other disruptions. Contribute to a culture of operational excellence through dashboards, playbooks, and on-call improvements. Lead complex technical projects from planning through successful deployment with clear stakeholder updates. Partner closely with peers across database, cloud, and infrastructure engineering teams as well as project management to investigate incidents and develop long-term roadmaps. Mentor junior engineers and foster a collaborative team environment. We’re looking for someone with 5+ years of professional software development experience building, deploying, and operating multi-tenant cloud services with a focus on operational excellence. Experience with large backend/compiled codebases, such as Rust or C/C++. Experience with containerization and orchestration platforms (e.g. Kubernetes). Experience with observability tooling (e.g. time series metrics, dashboards). Experience with distributed systems
About OfficeSpace: OfficeSpace Software provides the leading AI operating system for the built world, that helps teams plan, connect, and perform in the workplace. As a performance-based, PE-backed company, we hire based on merit and a willingness to do what it takes to succeed long-term. You’re a great fit for the role if you’re entrepreneurial, passionate, motivated by building at light speed, and an Agentic AI early adopter. Our world-class teams operate in the US, Canada, and Costa Rica in a culture of trust, respect, growth, and impact. About the role As a Lead Software Engineer at OfficeSpace, you'll shape the future of our workplace platform by building scalable products, leading technical execution, and raising the engineering bar across the team. You'll combine deep technical expertise with strong engineering leadership. From architecting distributed systems to mentoring teammates and designing AI-powered engineering workflows, you'll help us deliver enterprise software that is reliable, secure, and built to scale. This is a hands-on leadership role. We provide the platform. You drive the impact. What you'll do - Lead the architecture, design, and delivery of high-performance applications using Ruby on Rails, React, and modern cloud technologies. - Build scalable, maintainable software that supports enterprise customers across a rapidly growing platform. - Design AI-assisted engineering tools, workflows, and automation that improve developer productivity, code quality, and customer outcomes. - Own projects end-to-end—from technical discovery through production deployment and long-term ownership. - Shift quality left by embedding automated testing, code quality practices, and continuous validation early in the development lifecycle. - Drive predictable delivery by managing scope, balancing technical debt, protecting sprint commitments, and reducing reactive work. - Establish performance benchmarks and continuously optimize application spe
You’ll shape the future of a business‑critical platform as the technical lead across both product engineering and cloud infrastructure. You’ll modernize a mature .NET application running on AWS today, while steering its evolution toward a cloud‑native, React/Node.js, AI‑enabled architecture. If you enjoy owning architecture end‑to‑end, from backend and frontend through CI/CD, DevOps, and AWS infrastructure, this role gives you real influence at Staff Engineer level and the opportunity to set engineering standards that others follow. You’ll spend your time leading complex .NET and React features, designing scalable AWS infrastructure with Infrastructure as Code, and building automation that makes releases fast, safe, and repeatable. You’ll work on performance, reliability, and modernization in equal measure—fixing what’s slowing the platform down today and designing what it will look like in the next generation. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture and development of enterprise .NET services and APIs that power a business‑critical platform. Design and operate AWS infrastructure (using AWS CDK in TypeScript) to support secure, scalable, multi‑environment deployments. Build and optimize CI/CD pipelines (AWS CodePipeline, CodeBuild, Windows build agents) to make shipping .NET and React changes fast and reliable. Drive modernization initiatives across the stack, including clean architecture, refactoring legacy components, and reducing technical debt. Design and tune PostgreSQL and MSSQL database solutions for performance, scalability, and reliability. Mentor engineers and influence engineering practices across teams, raising the bar on cloud, DevOps, and software design. These are the essentials you’ll need to get an interview Significant experience (typically 8+ years) delivering and operating scalable enterprise software, owning both application code and cloud infrastructure. Deep hands‑on expertise with C
Opportunity Overview: We are seeking a Senior Data Engineer to contribute to the design and delivery of our cloud-native healthcare data platform. You will implement scalable data solutions built on AWS, Apache Iceberg, Lake Formation, Glue Catalog, Athena, dbt, and modern orchestration frameworks. This role combines strong hands-on engineering with collaboration across platform, analytics, and business teams. What You'll Do Data Engineering Delivery Deliver complex data engineering projects in collaboration with cross-functional teams Drive technical execution from design through production deployment Implement scalable data patterns and reusable frameworks Design and implement batch and near-real-time pipelines Build reusable ingestion, transformation, validation, and publishing frameworks Support modernization of legacy workloads Contribute to Apache Iceberg implementation and optimization Apply standards for schema evolution, partitioning, compaction, and metadata management Ensure efficient storage and query performance Implement data quality frameworks and validation layers Support observability and monitoring practices Contribute to operational excellence and reliability improvements Participate in architecture and design discussions Conduct and participate in code reviews Mentor junior engineers and share best practices ISMS roles and responsibilities Good knowledge of Information security Oversee specific business processes within the ISMS. Responsible to manage the ISMS documentation, conduct risk assessments, and implement risk treatment plans. Risk Owners are responsible for identifying, assessing, and managing risks within their areas of responsibility. They are also responsible for implementing risk treatment plans. Conduct the BCP and other test related to information security continuity along with CISO Responsible for monitoring and reporting on the performance of the ISMS. Responsible for implementation of security policies and procedures and report
Opportunity Overview: We are seeking a Lead Data Engineer to drive the design and delivery of our cloud-native healthcare data platform. You will lead the implementation of scalable data solutions built on AWS, Apache Iceberg, Lake Formation, Glue Catalog, Athena, dbt, and modern orchestration frameworks. This role combines deep hands-on engineering with technical leadership and collaboration across platform, analytics, and business teams. What you’ll do: Lead Data Engineering Initiatives Lead delivery of complex data engineering projects across multiple teams Drive technical execution from design through production deployment Establish scalable implementation patterns Build and Optimize Data Platforms Design and implement batch and near-real-time pipelines Build reusable ingestion, transformation, validation, and publishing frameworks Support modernization of legacy workloads Lakehouse Engineering Lead Apache Iceberg implementation and optimization Define standards for schema evolution, partitioning, compaction, and metadata management Ensure efficient storage and query performance Data Quality and Reliability Implement data quality frameworks Drive observability and monitoring practices Improve operational excellence and reliability Technical Leadership Review architecture and design proposals Conduct code reviews and engineering reviews Mentor engineers and establish best practices ISMS roles and responsibilities: Good knowledge of Information practices. Assist the manager in all the information security activities implementation and maintenance process. Ensuring the team and imparted with Competence related to Information security Responsible for implementation of security policies and procedures and report any issues to the Information Security Manager. Required Qualifications: 8–12 years of Data Engineering experience. Experience leading enterprise-scale data initia
Get new cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime