Datadog is seeking a Director of Product Management to lead our AI Observability portfolio and shape how organizations build, monitor, and scale AI systems in production. This role leads LLM Observability and helps define the next wave of innovation across GPU Monitoring, Distributed AI Monitoring, and emerging research-oriented tooling such as Model Lab. You will set the vision and strategy for this rapidly growing area, expanding established products while incubating new capabilities that deliver deep visibility into AI infrastructure, model performance, and distributed AI environments. As AI becomes core to modern applications, this team plays a critical role in ensuring customers can deploy and scale AI with confidence. We’re looking for a builder-minded product leader with strong technical depth and hands-on curiosity - someone who has built or worked closely with AI-powered products and understands the realities of production AI. You will lead a team of product managers and partner closely with engineering and design to advance Datadog’s leadership in AI observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the vision and strategy for AI-driven products, ensuring alignment with overall company goals and customer needs. This will include managing our embed program to enhance the capabilities of existing products as well as developing dedicated and independent AI products. Lead and mentor a team of product managers, helping them grow and advance their careers while ensuring the delivery of high-quality, AI-powered features. Collaborate with cross-functional teams including engineering, data science, marketing, and sales to deliver AI product solutions that meet customer needs and business objectives. Identify new opportunities for
Jobiba hiring network
Production Operator Jobs
3,235 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current production operator jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Applied AI is where Datadog's ambitious AI bets get built and shipped ( Bits Chat , updog ). We sit at the intersection of research and product: turning promising capabilities from Datadog AI Research lab and the research community into production systems that reach real customers. The team builds specialized models that replace frontier models where they are not necessary, making AI capabilities faster, cheaper, and more secure. The mandate is to move fast from idea to customer impact, and when a product finds its footing, to set it up for growth. As a Manager I in Applied AI, you will lead a team of engineers and applied scientists working on one of these challenges. You will define technical direction, run short feedback loops, make deliberate decisions about what to pursue or stop, and work closely with product managers, research teams, and cross-functional partners to ship AI capabilities that matter. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Lead and develop a team of engineers and applied scientists focused on cost-efficient specialized models and AI security capabilities Work closely with product managers, research teams, and cross-functional partners to shape the team's bets from initial framing through to broader adoption, with a clear definition of success criteria at each stage Own end-to-end delivery of high-quality AI systems, from early research exploration to production-grade reliability, with high standards for operational excellence, system reliability, and technical quality Navigate the unique challenges of shipping AI-powered products: balancing quality, latency, cost, and safety considerations. Drive evaluation and iteration practices for AI systems: define the quality bar and guide the team in building the offline
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based out of our Dublin or Cork office or remotely in Ireland. What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA teams. Success in this role means smoother launches, clearer roadmaps, stronger reliability metrics and an SRE organization that's better-equipped to deliver predictability at scale. This role can be based remotely on the East Coast What You'll Do Drive Program Planning & Execution – Define program scope, milestones, and success criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep work clearly tracked in Jira, and deliver on time Strengthen Production Reliability – Lead change management and launch readiness programs. Partner with SREs and product teams to define and operationalize SLOs/SLIs, and use incident data, metrics, and capacity signals to drive prioritization and continuous improvement Lead Cross-Functional Coordination – Align SRE with Security, Compliance, Cloud platform, and other engineering teams. Coordinate cross-team incident response, ensure clear follow-through, and build trust as the go-to driver of complex, multi-team efforts Build Scalable Systems & Processes – Design lightweight frameworks and communication patterns that help SRE deliver reliably at scale. Work yourself out of the "hero" role by leaving teams better-equipped to execute independently Requirements 8+ years in technical program management, engineering management, or a comparable technical role partnering with software engineering teams Proven track record leading large-scale, cross-team platform initiatives through ambiguity and change Strong knowledge of production change management, software development lifecycle, and reliability metrics (SLOs, SLIs) Skilled at shaping roadmaps and managing dependencies Able to query and interpret metrics, logs, or other data s
Atlas Growth is a cloud engineering group whose mission is to guide customers through their app development journey—from cluster configuration, data modeling and load testing, to running a production workload at scale. We use an in-house experiments platform which helps us validate our features quickly, releasing only the work that positively impacts our customers. Our engineers participate in cross-functional “squads” with product, design, analytics, and research focusing on a single metric (e.g. retention). Our engineering team is part of a larger Atlas Core Engineering org, building foundational elements of MongoDB’s developer data platform. Atlas Growth 2 builds customer-facing features in Atlas and sits alongside other Growth engineering teams. Recent projects include an AI Chatbot for cluster creation, a recommendation system that offers tips for better database performance, and a pricing page designed to optimize conversion rates. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Role Overview Atlas Growth seeks a mid-level software engineer (Software Engineer 3). SE3s are solid contributors to projects they work on and often lead projects of their own. They act in accordance with MongoDB’s core values and leadership principles, and are actively working toward a Senior role. Candidate Profile 3+ years of software engineering experience, with fluency in TypeScript/JavaScript, and experience with a modern framework (e.g. React) Proficiency in Java, Go, C++/C, or a similar compiled language is a plus Experience writing database queries, either document-based or relational Experience writing and reviewing technical specs, and leading small projects Interest in A/B testing or product design Expectations Contribute readable and well-tested code to ongoing projects Collaborate closely with product and design partners to implement and iterate on new customer facing features Write scope and technical spec docs for new projects
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Has experience working with AI coding agents and can demonstrate building high quality context to yield high quality outputs Has experience designing and implementing medium-to-large software projects, including driving design reviews and mentoring less-senior engineers Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Str
Senior Forward Deployed Engineers (FDE) sit at the intersection of enterprise customer environments and a fast-moving internal product. They partner directly with customers to design, build, troubleshoot, and improve production solutions, and they are ultimately accountable for helping customers get to production. This role is best suited for engineers who want to own technical outcomes end to end: understanding what a customer is trying to build, shipping the integration, and feeding learnings back into the product roadmap. This role will be based remotely in the United States (East Coast). Responsibilities Customer success Serve as the primary technical owner for customer engagements from initial discovery through production rollout Understand each customer's architecture, constraints, and definition of success, and drive toward that outcome Manage expectations, communicate risks clearly, and help customers navigate technical decisions with confidence Technical integration Build the connectors, pipelines, and supporting tooling needed to make the platform work inside real enterprise environments Write production-quality code, troubleshoot issues, and implement fixes directly in active workstreams Work effectively within customer environments that have different stacks, infrastructure, and integration constraints Product feedback loop Capture product feedback with precision, including logs, reproduction steps, and a clear proposed path forward Use customer engagements to identify product gaps, surface recurring patterns, and help improve the product roadmap Document technical decisions and tradeoffs clearly so product and engineering teams can extend the work Workstream collaboration Partner closely with product and engineering teams to help turn field patterns into reusable product capabilities Contribute directly in focused workstreams by helping drive technical design, implementation, and delivery Make sound engine
Senior Forward Deployed Engineers (FDE) sit at the intersection of enterprise customer environments and a fast-moving internal product. They partner directly with customers to design, build, troubleshoot, and improve production solutions, and they are ultimately accountable for helping customers get to production. This role is best suited for engineers who want to own technical outcomes end to end: understanding what a customer is trying to build, shipping the integration, and feeding learnings back into the product roadmap. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Responsibilities Customer success Serve as the primary technical owner for customer engagements from initial discovery through production rollout Understand each customer's architecture, constraints, and definition of success, and drive toward that outcome Manage expectations, communicate risks clearly, and help customers navigate technical decisions with confidence Technical integration Build the connectors, pipelines, and supporting tooling needed to make the platform work inside real enterprise environments Write production-quality code, troubleshoot issues, and implement fixes directly in active workstreams Work effectively within customer environments that have different stacks, infrastructure, and integration constraints Product feedback loop Capture product feedback with precision, including logs, reproduction steps, and a clear proposed path forward Use customer engagements to identify product gaps, surface recurring patterns, and help improve the product roadmap Document technical decisions and tradeoffs clearly so product and engineering teams can extend the work Workstream collaboration Partner closely with product and engineering teams to help turn field patterns into reusable product capabilities Contribute directly in focused workstreams by helping drive technical design, implementation, and delivery Make sound engineering decisions under
Come join and lead the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for a talented Lead Engineer to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration, and lives MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. If you want to lead a fast-growing team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 3+ years of experience managing a team of software engineers, including hiring, performance and growth management, compensation planning, and mentoring You have 8+ years of experience building production-quality systems software with large backend/compiled codebases, ideally in Rust. Bonus points for experience with performance profiling, network protocols, TLS, and connection management You have strong technical judgment that you use to effectively guide engineering decisions in security-sensitive or networking-adjacent domains You put the customer first and don't hesitate to cross team boundaries in search of the right solution Solid experience in designing, writing, testing, maintaining, and operating mission-critical software systems Bonus points Professional or advanced academic expertise in the domains of security or networking You enjoy coaching, career development, and creating growth opportunities to help your
Cloud Operations Engineers are responsible for building internal tools and process automation. Day-to-day duties are creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases, and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. We are looking for a Lead with strong technical leadership experience as well as technical depth who is looking to collaborate closely with Cloud Operations Engineering Management in building and maintaining a high-performing team that delivers high quality outcomes while fostering psychological safety and professional growth. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Core responsibilities Team leadership: partner with and assist COE Management with the tasks of providing ongoing technical feedback to engineers, support their growth and creating an inclusive team environment Execution and delivery: play a key role in guiding team members through project deliverables ensuring high quality outcomes while also assisting in meeting or resetting timelines when required Time management: between assisting team members with day to day tasks ranging from incident to project management Cross-functional collaboration: work closely with Product, Technical Services and R&D to surface team’s pain points and drive alignment with the goal of providing an excellent user experience to the end customer Coordinate with Lead counterparts within Cloud Operations as well as Technical Services to ensure our uptime guarantees to the MongoDB Atlas customer base Assist and collaborate with the team on scoping, designing, deploying and ongoing maintenance of systems that focus on reducing mean time to resolve customer incidents Detec
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 provides an unparalleled authentication experience for hundreds of millions of users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in Europe to ensure our production systems are not only operational but also resilient, scalable, and ready for exponential growth. This isn't just about keeping the lights on; it's about directly contributing to the platform's core resiliency and robustness. You'll be a hands-on builder, crafting solutions that make our system more reliable by design. What you’ll do: Design and build custom software in Go to enhance the platform's reliability, resiliency, and redundancy. Partner with engineering teams to embed reliability principles, improving the availability, performance, and observability of our services. Use your deep understanding of infrastructure and observability principles to identify opportunities for improvement within the product and implement solutions. Contribute to our follow-the-sun on-call rotation, providing rapid, effective response to critical incidents and using your expertise to troubleshoot, mitigate or accurately escalate production issues. Because our team is globally distributed, your on-call shifts will only occur during your standard local working hours. Develop and refine our SRE tooling and proc
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t
Get new production operator jobs by email
Daily job updates · Unsubscribe anytime