Jobiba hiring network

Devops Engineer Observability Manager Consultant Jobs

568 active opportunities · Updated for September 2026

Market range: $127K – $174K/yr

Fresh results

15 shown

Explore current devops engineer observability manager consultant jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
Okta
📍 Bengaluru, IndiaFull-time$127K – $174K/yr · Jobiba est.
29 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from

pythonsqlpostgresql
View job →
O
Okta
📍 IndiaFull-time$127K – $174K/yr · Jobiba est.
29 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
O
Okta
📍 IndiaFull-time$127K – $174K/yr · Jobiba est.
29 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
O
Okta
📍 IndiaFull-time$127K – $174K/yr · Jobiba est.
29 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
P
Postman
📍 San FranciscoFull-time$3.1M – $3.3M/yr
1mo ago

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo

aigorust
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Account Executive - Enterprise Sales - Germany Berlin, Germany OR Munich, Germany We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. We are looking for an ambitious Account Executive to join our growing German team. If you are a results-oriented sales professional ready to help enterprise customers optimize their digital footprint, we want to hear from you. Your Opportunity As an Enterprise Account Executive, you will be responsible for identifying, qualifying, and closing new business within a defined geographic or vertical territory. You will manage the end-to-end sales cycle, from initial prospecting to contract negotiation. You will work closely with our solution engineers and marketing teams to demonstrate the value of New Relic and displace competitors in the observability space. What You’ll Do Pipeline Generation: Proactively prospect into a list of target enterprise accounts. You will own your "top of funnel" by combining inbound leads with your own strategic outbound efforts. Sales Execution: Manage a disciplined sales process to meet and exceed monthly and quarterly quotas. You will lead discovery calls, coordinate product demos, and navigate the procurement process. Customer-Centric Selling: Understand the specific pain points of German enterprise customers

gitrestai
View job →

About Datadog: We're on a mission to build the best platform in the world to defend the enterprise from code-to-cloud-to-runtime. Used by thousands of companies globally, Datadog security products uniquely leverage Datadog's unified security and observability platform so Security, DevOps and SRE can collaborate rapidly and seamlessly to deliver better detection, prioritization and remediation. Our product and engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity: In this competitive market, the Director of Product Management for Cloud Security and Platform will play a mission-critical role in providing product and strategy leadership to grow Datadog's market share through differentiation, innovation and compelling customer value. This leader will work with a talented and growing team of product managers and engineers to build and grow multiple product lines that play an essential role for our customers' cloud security programs and the shared platform capabilities supporting all Datadog security products, to establish Datadog as a security industry leader. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Run and grow multiple product lines to meet revenue and business targets with the goal of building a multi-hundred million dollar annual business. Lead and own product strategy and roadmap for accountable security product lines, fully aligned to revenue and business goals and with compelling differentiation and customer value. Ensure predictable roadmap execution across direct and partner teams to achieve product and business outcomes required to meet the revenue and business goals. Analyze and develop pricing and packaging strategies to maximize revenue through attaching deep understand

aigorust
View job →
T
Twilio
📍 - IrelandFull-timeRemote$127K – $174K/yr · Jobiba est.
29 days ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Software Engineer on Twilio’s platform engineering observability team. About the job This position is needed to help our platform engineering observability team. Twilio is undergoing a large-scale observability transformation—and you can help shape the foundation. Observability is a strategic pillar and a key enabler for faster incident response, deeper customer-centric insights, and more cost-effective platform operations. As a Software Engineer on the Platform Observability team, you’ll play a critical role in re-architecting how telemetry flows and is utilized through Twilio—making it structured, accessible, affordable, and actionable. Over the next 3 years, Twilio is rebuilding nearly every component of our observability platform, from data collection to real-time analytics. You will drive core initiatives that shift Twilio from fragmented tooling and wasteful data sprawl to a unified, OpenTelemetry-first observabil

REMOTEpythonjavaaws
View job →
O
Okta
📍 WashingtonFull-timeFrom $147K/yr
29 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Auth0 Platform Observability team owns the observability tooling that monitors the Auth0 Platform, and we are looking for an Observability Engineer to help ensure that our Product and Platform Engineers can monitor and observe our platform while continuing to rapidly ship software that our customers love. Our engineers maintain and automate observability tooling for our entire platform, including metrics, logs, and traces. We are looking for engineers passionate about monitoring, observing, measuring uptime and availability, and ensuring platform stability. If you have experience within the Site Reliability Engineering (SRE) field or working as a Development Operations (DevOps) engineer, and you have a passion for Observability tooling, this position will allow you to further your learning and development in these areas. As a Senior Engineer on this team, you will act as a core technical leader. You will work cross-functionally to help integrate services with our instrumentation libraries, support product teams, and actively investigate incidents to identify our observability gaps. Responsibilities: Proven ability to champion observability best practices, acting as an educator who can effectively correct anti-patterns and teach other engineering teams how to build robust, standardized instrumentation. Be an expert in running services in production environments Contribute to the process of designing services for high growth and high availability. Provisi

node.jsawsazure
View job →
D
Datadog
📍 CaliforniaFull-timeRemote
28 days ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s

REMOTEpythonnode.jsai
View job →
N
Newrelic
📍 SpainFull-time
29 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand, and tackle issues, analyze performance, and get the most out of their software and infrastructure. This is a unique opportunity to shape the future of observability while pioneering the next generation of engineering: we are actively revolutionizing how we build software by embedding modern, agent-powered workflows directly into our daily development lifecycle. You’ll tackle complex distributed systems problems while helping drive internal innovation on the frontlines of AI-assisted engineering. About the team This position is for the Service Architecture Intelligence team. You will be building, improving, and maintaining a distributed service architecture capable of ingesting large volumes of data, analyzing spans and traces, and publishing them downstream so the UI can offer an exceptional experience to our customers. We work with data at scale, and our pipeline is built with a diverse tech stack (Java, Kafka, Redis, public cloud services, and more). You will work alongside a team of talented engineers solving complex distributed systems challenges. If you're passionate about performance and scale, and want to contribute to one of the largest and fastest-growing observability platforms while co-crea

javaredisgit
View job →
N
Newrelic
📍 SpainFull-time
29 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity As a Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with a proven track record of building and scaling resilient infrastructure. What you'll do Platform Orchestration: Work on a large-scale K8s infrastructure platform, ensuring high availability and performance. Automation: Drive the evolution of internal tooling to streamline platform delivery. This role requires Experience: Solid hands-on background in DevOps, Site Reliability, or Infrastructure Engineering. Kubernetes Mastery: Deep internal knowledge of K8s primitives (Deployments, StatefulSets, Services) and hands-on experience writing custom Kubernetes Operators. Golang Proficiency: Proficiency in Go, specifically for infrastructure automation and systems programming. Operations-Heavy Mindset: A proven track record of managing production environments and handling high-severity incidents. Cloud Infrastructure: Hands-on experience with cloud-native scaling tools (e.g., Karpenter, Cluster API) and Day 1/Day 2 operations of K8s clusters. Tooling: Familiarity with Helm and GitOps workflows (e.g., ArgoCD or Flux). Please note that visa sponsorship is not available for this position. Fostering a diverse, welcoming and inclusive environment is important to us. We work hard to make everyone feel comfortable bringing their best, most authentic selves to work every day. We cele

kubernetesgitrest
View job →
N
29 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand, and tackle issues, analyze performance, and get the most out of their software and infrastructure. This is a great opportunity to shape the future of observability, work on complex and interesting problems, and empower companies to thrive in an AI-first world. This position is for the Distributed Tracing Pipeline team. You will be building, improving, and maintaining a distributed service architecture capable of ingesting large volumes of data, analyzing spans and traces, and publishing them downstream so the UI can query this data to offer an exceptional experience to our customers. We work with data at scale, and our pipeline is built with a diverse tech stack (Kotlin, Java, Kafka, Redis, public cloud services, and more). You will work with a team of talented engineers, solving complex distributed system challenges and delivering solutions for our customers. If you're passionate about scale and want to contribute to one of the largest and fastest-growing observability platforms in the market, we want to hear from you. We are a fast-growing software company that cares about our product and culture. What you'll do Build, maintain, and scale backend services and their support tools Particip

javaredisgit
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto

awsazurekubernetes
View job →
N
Newrelic
📍 IrelandFull-time
29 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity As a New Relic Support Engineer, you know more about our products than any other function and you feel a sense of pride and satisfaction in helping customers through their never-seen-before technical issues. We are serious about keeping our skills sharp so that we can provide outstanding assistance in a constantly evolving technical landscape. We emphasize training, knowledge and customer empathy, your learning opportunities will never end. Are you ready to become our next Support Engineer passionate about today’s Infrastructure issues? What you'll do Collaborate across teams to assist in solving complex technical customer problems across our product suite Strong analytical and technical troubleshooting skills Impeccable customer service skills and display genuine empathy towards customers. Work closely with our software engineering teams to resolve advanced customer issues. Support New Relic customers by resolving various installation, configuration, and data exploration requests. Be an advocate for our customers to our Product Organization by providing feedback on feature requests and bugs that improve the customer experience of the New Relic platform. Advance your skills through additional training and exposure to other features and capabilities of our Products This role requires Has a track record of providing superior end-user software support via email, phone, and social media Love delighting customers, even those who are having a tough day! Experience b

javascriptjavaaws
View job →
🔔

Get new devops engineer observability manager consultant jobs by email

Daily job updates · Unsubscribe anytime