As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i
Jobiba hiring network
Cloud Operations Engineer Jobs
2,288 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
As a Staff Engineer on the Data Platform Experience team, you'll help shape how Datadog engineering teams build, operate, and evolve products on the Observability Data Platform. You'll lead the design and delivery of shared platform capabilities that reduce developer friction, improve operational visibility, and enable engineering teams to move faster with confidence. This role combines deep distributed systems expertise with technical leadership across multiple teams, influencing platform strategy while remaining hands-on in the code. You'll have the opportunity to solve company-wide challenges spanning cost intelligence, operational tooling, platform health, and developer experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead strategic engineering initiatives that improve how product teams build, operate, and evolve services on the Observability Data Platform. Design and build scalable platform capabilities for cost intelligence, including cloud cost allocation, trend analysis, and optimization recommendations. Develop operational intelligence and self-service tooling that helps engineering teams understand platform health, troubleshoot incidents, and improve operational efficiency. Drive reusable platform services and developer workflows that increase engineering autonomy while reducing operational complexity across multiple products. Provide technical leadership across teams by influencing architecture, mentoring engineers, and raising engineering standards through hands-on technical contributions. Participate in the team's on-call rotation and continuously improve platform reliability, observability, and operational excellence. Who You Are: You have experience designing and building large-scale SaaS or cloud platforms with deep expertise i
We’re looking for an Engineering Manager to lead our Sensitive Data Scanner (SDS) Telemetry team. The SDS group’s mission is to be the world’s easiest-to-use tool to discover, classify, manage, and report sensitive data risks across cloud, on-premise, and code environments. This team builds and scales the detection capabilities that scan all telemetry data flowing into Datadog — logs, APM spans, and RUM events — operating in streaming, at processing time, and at very large scale. You’ll lead a small, close-knit team based in Paris, with the opportunity to shape how the team grows as SDS Telemetry’s scope expands. It’s a chance to combine hands-on technical leadership with direct customer and product impact in the security and observability space. At Datadog, we place value in our office culture — the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead and grow a team of engineers building real-time sensitive data detection across Datadog’s Logs, APM, and RUM telemetry pipelines Partner closely with the Logs, APM, and RUM teams, plus Datadog’s Trust & Safety team, to align on roadmap and integration priorities Shape product direction by working closely with Product, grounding decisions in customer needs and business impact Stay hands-on: contribute to design decisions and participate in the team’s on-call rotation Recruit, mentor, and develop engineers as the team grows beyond its initial size Help build a strong engineering culture as part of Datadog’s broader Sensitive Data Scanner group Who You Are: You have experience building and shipping revenue-generating products, with strong product acumen and a customer-first mindset You have hands-on experience with Go and/or Java, and a track record building distributed, streaming systems at scale You have experience managing engineers — or are
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Has experience working with AI coding agents and can demonstrate building high quality context to yield high quality outputs Has experience designing and implementing medium-to-large software projects, including driving design reviews and mentoring less-senior engineers Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Str
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers and Architecture teams, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term performance, reliability and scalability goals in a platform with an exponential growth plan for the coming years. With Okta’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical enterprise and consumer applications. Skills Exceptional communication skills, including technical writing in the English language Systematic problem-solving approach, coupled with a strong sense of ownership and drive Understanding of microservices, cloud infrastructure (AWS, Azure), databases (SQL, No-SQL, Key/Value), containers (docker, kubernetes), web technologies (web sockets, http) and networking (SSL, routing, VPN) Live and breathe SLIs, SLOs, error budgets and SLAs Strong belief in automating everything and reducing toil for yourself and teammates Loves to work as a team, but is able to work effectively in a remote environment where tasks may be self-driven Knowledge
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology, anywhere, on any device or app. Our flexible and neutral products, Okta Platform and Auth0 Platform, provide secure access, authentication, and automation, placing identity at the core of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box - we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. About Okta’s Enterprise Access Team Okta is The World’s Identity Company. We free everyone to safely use any technology—anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. The Enterprise Access team drives billions of authentications every month. The team builds and supports single sign-on, strong authentication, provisioning, and threat protection technologies. Our Enterprise Access service runs in the cloud on a secure, reliable, extensively audited platform with 99.99% availability. About the role We’re looking for a Principal Software Engineer for the SIW Platform team. Operating under the larger Enter
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Team Okta's Core Engineering team is responsible for building and evolving shared infrastructure and services that lay the foundation for what other engineering teams build on. We're in charge of common shared services like distributed cache, configuration management, frameworks for async job management, internal tooling for developer support, and email pipeline, to name a few. We're cloud native, where redundancy, multi-tenancy, scale, resource optimization and resiliency are first class citizens. With Okta's mantra of 'Always On!' there's never a dull moment. Our biggest asset is our team of passionate engineers and technically minded managers. Role: This is an opportunity for an experienced Backend engineer to join our growing Core Platform team based out of Bengaluru. In this role, you will get to work with highly skilled and talented engineers throughout the organization to build and manage some of the critical platform services powering Okta’s products and infrastructure. This role requires a blend of high-level architectural thinking and hands-on execution to build resilient, high-performance backend services. We are not only passionate about building services but operating them at scale, making them resilient to provide a seamless service to our customers. You'll be leading a team of highly skilled and talented team players who're proud of what they own and deliver. Our elite team is fast, creative and flexible; with a weekly release
Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Capacity & Efficiency Engineering team builds the software that manages, governs, and reduces Robinhood's AWS cloud spend, operating at the intersection of cloud infrastructure, data engineering, and FinOps. The team owns the full lifecycle of cloud cost: the data platforms that make spend transparent and attributable, the anomaly detection and forecasting systems that make it predictable, the automation that continuously rightsizes infrastructure at fleet scale, and the capacity planning and commitment strategy that keep a bursty, latency-sensitive trading platform both reliable and cost-effective. Our work carries CEO-level visibility and has already driven millions of dollars in annualized savings. We partner closely with Data Science, Infrastructure, Finance, and product
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the Onboarding team, you'll build the systems that power user onboarding and KYC across every Coinbase product and market. This team enables Coinbase to launch in new countries and ship new products quickly while meeting compliance and regulatory requirements. You'll own backend services that drive onboarding conversion at global scale and shape the architecture that makes rapid expansion possible. What you'll do: Own the design, development, and scaling of backend services in Golang that power onboarding and KYC workflows across multiple products and geographies. Build scalable, service-oriented systems using modern cloud infrastructure and industry best practices to solve novel platform challenges. Drive improvements to onboarding and MTU conversion funnels by identifying bottlenecks and shipping data-informed solutions. Partner with engineers, product managers, designers, and senior leadership to translate product and technical vision into quarterly execution roadmaps. Lead cross-cutting design reviews within your product area, ensuring security, operational integrity, and architectural clarity across all features shipped. Required Skills and Experience: 5+ years of professional software engineering experience building, scaling, and maintaining production backend services in a service-oriented architecture. Proven track record desi
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our transport network serves the needs of millions of people every day who want to get from one place to another using Lyft cars, bikes and scooters, with public transportation, or on foot in the most efficient way. To serve these needs, we need to suggest the fastest, most affordable and safest routes. We achieve this by processing millions of rides, taking into account the latest traffic information and analyzing the preferences of drivers. To strengthen our efforts, we are hiring a Software Engineer who will work on improving our routing engine, on building cloud-based services that can process millions of route requests per day and on creating workflows to process and analyze the data collected from our rides. For this we are looking for someone who has a strong background in software architecture and algorithms and understands at the same time how to build efficient data processing pipelines. Our technology stack is based on the latest technologies such as AWS, Kubernetes, Apache Airflow, Flink, and Kafka. We use AI and machine learning to improve development productivity, optimize customer support workflows, and enhance operational efficiency — all without compromising code quality or security. You will work with incredibly passionate and talented colleagues from software engineering, machine learning and data science on projects that delight millions of passengers and drivers. Responsibilities: Drive high-impact projects and innovate new solutions to provide the best routing experience possible Build and deploy mission critical algorithms and services that can serve millions of requests per day (backend and data-streaming services written in C++, Go, and Python) Analyze rides, understand customer pain points, prioritize and size feature requests and help design solutions based on o
About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist
About the team The Applied AI Engineering (AAE) team is responsible for helping developers and enterprises turn the potential of generative AI into real-world impact. We act as trusted advisors and technical partners to customers and ecosystem partners, helping identify high-impact AI use cases and bring them into production through strong architectural guidance and hands-on execution. The Partner Applied AI Engineering organization works closely with strategic cloud providers, systems integrators, consultancies, and implementation partners to scale successful adoption of OpenAI technologies. As the leader of the AWS Partner AAE pod, you will manage a team of Applied AI Engineers focused on enabling AWS-aligned partners and their customers to build, deploy, and operationalize AI applications on OpenAI’s platform. About the role We are seeking a Manager, Partner Applied AI Engineering – AWS to lead a team of Applied AI Engineers supporting strategic AWS ecosystem partnerships. In this role, you will own the technical success strategy for AWS-aligned partners and help build scalable, repeatable ways for partners and their customers to adopt OpenAI technologies. Your team will guide partners and customers across the full AI implementation lifecycle—from identifying and shaping high-value use cases to solution design, architecture, production deployment, optimization, and adoption growth. You will work cross-functionally with internal and external stakeholders across Sales, Partnerships, Product, Research, and Engineering to ensure the voice of partners and customers informs our platform roadmap and how we bring OpenAI technology into production at scale. This role requires a blend of technical depth, customer leadership, operational rigor, and people management. Success will be measured through production deployments, partner technical maturity, API adoption growth, team development, and the overall impact of the AWS partner ecosystem. This role is based in our San Fra
About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We’re seeking an exceptional Staff - Principal level offensive security domain expert to build agents that continuously identify and coordinate remediation of vulnerabilities across OpenAI’s infrastructure and applications. You will be the technical owner of this effort, combining deep offensive security judgment with agent engineering to build a production system that can operate safely and reliably at scale. As OpenAI increasingly uses automation throughout the company, we believe our security testing must become increasingly automated as well. Advances in model capabilities create an opportunity to test more of our attack surface than would be possible through human effort alone and a need to ensure that we remain ahead of those same capabilities as they become available to attackers. In this role, you’ll build a portfolio of specialized agents that develop a deep understanding of OpenAI’s infrastructure, applications, processes, and security boundaries. These agents will combine internal context with feedback from running systems to explore our cloud environments, Kubernetes clusters, web applications, endpoints, external attack surface, and other high-value targets. The goal is for agents to not only discover vulnerabilities, but also to validate exploitability, document impact, drive remediation, and verify fixes. Success will be measured through outcomes like vulnerabilities fixed, attack surface covered, and performance on evals
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. About the Role As a Site Reliability Engineer at Ema, you will own the stability, availability, and operational health of our agentic AI platform across customer environments. You'll work closely with Engineering and DevOps to provision infrastructure, drive deployment excellence, and keep production running at the quality bar our enterprise customers expect — 99.9%+ uptime, proactive incident response, and continuous improvement. What You'll Do Infrastructure & Deployment Design and provision cloud infrastructure (GCP, Azure, AWS) tailored to customer environments, with security, scalability, and compliance built in Execute on-call SaaS deployments with minimal downtime; automate and optimize deployment workflows end-to-end Production Stability & Observability Monitor logs, alerts, and metrics to maintain SLA commitments and catch issues before they escalate Diagnose and resolve production incidents with speed and rigor; drive root cause analysis and permanent fixes Collaborate with DevOps to enhance monitoring dashboards and alerting frameworks; deliver clear system health reporting to internal and customer stakeholders Documentation & Knowledge Management Maintain de
Get new cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime