Jobiba hiring network

Lead Devops Engineer Observability Manager Manager Consultant Jobs

15 active opportunities · Updated for September 2026

Market range: $127K – $174K/yr

Fresh results

15 shown

Explore current lead devops engineer observability manager manager consultant jobs. Use filters to narrow by work mode, employment type, experience and date posted.

D
1mo ago

We are Datadog's in-house product experts. The Technical Solutions team enables Datadog’s worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. As a Technical Account Manager 2 (TAM 2), you’ll serve as a trusted advisor to our strategic customers, accelerating their adoption of the Datadog platform and enabling long-term success. TAM 2s bring deep technical expertise, refined customer skills, and consultative insight into how monitoring, observability, and DevOps practices translate to business value. At Datadog, we place value in our office culture—the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical advisor to 3 primary enterprise accounts, ensuring successful product adoption and effective usage of the Datadog platform. Lead enablement and adoption sessions across core product areas tailored to your customer’s architecture and business needs. Analyze customers’ IT Operations environments and workflows to recommend configuration, product usage, and performance improvements. Deliver technical business reviews, health checks, and account maturity assessments, contributing to customer QBRs with impactful recommendations. Escalate product issues appropriately and advocate for your customers’ needs with Datadog’s Product and Engineering teams. Create executive-level summaries and insights that tie platform usage to business outcomes. Participate in internal TAM strategy sessions and contribute to team learning through feature presentations, case studies, or best-practice sharing Who You Are: You have 2+ years of experience in a technical customer-facing role (TAM, Solutions Architect, SRE, DevOps Engineer, or similar) within the cloud or observability space. You’re confident with at least two public cloud platforms (e.g. AWS, Azure, GCP)

pythonawsazure
View job →

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Database Observability is a critical pillar of New Relic's platform strategy. We are looking for an experienced Engineering Manager (M3) to lead a senior, high-performing team building next-generation Database Observability products. Your team will own the entire lifecycle of critical telemetry data flows from lightweight database agents and high-throughput ingestion pipelines to intelligent DB recommendation engines and autonomous DB AI Agents. You will lead a team that includes senior and Lead-level engineers with deep domain expertise in distributed systems and AI. Your primary value will come from setting strategic technical direction, enabling their best work, and fostering a high-accountability culture while partnering closely with Product and Design to deliver features that directly drive New Relic's Database Observability. What you'll do Manage a full-stack engineering team (6–8 engineers) spanning backend systems, database telemetry, agent engineering, and UI workflows. Own end-to-end delivery sprint planning, roadmap execution, system quality, and operational excellence for critical database ingesti

pythonjavasql
View job →
R
Roblox
📍 San MateoFull-timeFrom $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Observability team builds the infrastructure that empowers engineers to understand, operate, and improve the Roblox platform and ecosystem. Our team owns the end-to-end observability stack across telemetry, distributed tracing, logging, profiling, storage systems, and developer-facing visualization tools. We are looking for an Engineering Manager to lead the next generation of AI-powered observability platforms. In this role, you will help build intelligent systems that leverage AI to revolutionize CI/CD, testing, and DevOps workflows — enabling engineers to move faster, improve reliability, and operate large-scale distributed systems with greater efficiency and confidence. This is a highly impactful leadership role at the center of Roblox infrastructure. Your work will directly improve developer productivity, platform reliability, and operational excellence across the company. You will partner closely with infrastructure, product engineering, and AI platform teams to shape the future of developer tooling and autonomous operations at scale. You Have 3+ years of engineering management experience with a proven track record of hiring, mentoring, and growing high-performing teams. Strong ex

awsci/cdgit
View job →
N
Newrelic
📍 IndiaFull-time
1mo ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity Are you passionate about building foundational technology that fuels the world’s digital innovation? At New Relic, we provide the leading unified data platform for all things observability – helping engineers, developers, and operators make decisions using data at every stage of the software lifecycle. Our New Relic Control team is at the heart of that mission, creating groundbreaking capabilities that orchestrate, manage, and optimize observability agents and telemetry pipelines at scale. We’re looking for a Senior Product Manager to lead a new strategic initiative within our Pipeline Control team. In this role, you’ll be responsible for defining vision, strategy, and roadmap for an enterprise-grade solution that orchestrates observability pipelines to streamline telemetry data in flight. You'll collaborate closely with engineering and product design to solve complex challenges around instrumentation, configuration, and data management with new systems and experiences. If you love turning ambitious ideas into impactful enterprise products that delight customers, let's talk! What You'll Do Define and drive the product strategy for our next-generation observability pipelines solution and align your vision with broader company goals Lead cross-functional collaboration with Engineering, Product Design, Sales, and Marketing to translate customer insights and technical opportunities into impactful product capabilities Evangelize the product vision internally

kubernetesgitlinux
View job →
GH
greenhouse,Cohere Health
📍 HyderabadFull-time₹2K – ₹2K/yr
11 hrs ago

Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I

D
Datadog
📍 New YorkFull-timeFrom $2.9M/yr
1mo ago

We're on a mission to build the best platform in the world to defend the enterprise from code-to-cloud-to-runtime. Used by thousands of companies globally, Datadog security products uniquely leverage Datadog’s unified security and observability platform so Security, DevOps and SRE can collaborate rapidly and seamlessly to deliver better detection, prioritization and remediation. Our product and engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. In this competitive market, the Group Product Manager for Code Security will play a mission-critical role in providing product and strategy leadership to grow Datadog’s market share through differentiation, innovation and compelling customer value. This leader will lead a talented and growing team of product managers and work with world class engineers to build and grow multiple Code Security products that play an essential role for our customers’ code security programs, and growing Datadog into a security industry leader. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Run and grow multiple Code Security products to meet revenue and business targets with the goal of building a multi-hundred million dollar annual business. Lead and own product strategy and roadmap for accountable security products, fully aligned to revenue and business goals and with compelling differentiation and customer value. Ensure predictable roadmap execution across direct and partner teams to achieve product and business outcomes required to meet the revenue and business goals. Analyze and develop pricing and packaging strategies to maximize revenue through attaching deep understanding of market dynamics and other strategic leverage points. Drive GTM strategy with GTM partner teams

aigorust
View job →
N
15 days ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. Service Levels is one of New Relic's most commercially critical products — it powers SLO compliance and reliability measurement for thousands of customers. We're looking for an experienced Engineering Manager to lead a senior, high-performing team building the next generation of service level management at scale. You'll lead a team that includes lead-level engineers with deep domain expertise, and your value will come from enabling their best work — not directing it. You'll own delivery, quality, and team health while partnering closely with product and design to ship features that directly impact New Relic's commercial momentum. What you'll do Lead a full-stack engineering team of 6-8 across backend (Java/Spring Boot) and frontend (React/TypeScript) Own end-to-end delivery — sprint planning, execution, quality, and release Set clear expectations, manage performance equitably, and develop engineers at every level Partner with Product Manager and XD to translate requirements into technically sound, deliverable plans Drive architectural discussions and hold the team to engineering excellence standards Identify

typescriptjavareact
View job →
N
Newrelic
📍 IndiaFull-time
1mo ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. The Infrastructure product organization develops New Relic infrastructure instrumentation agents, next generation data processing and management services, vulnerability management, and security testing capabilities for on-prem and cloud customers. We work with data at a scale using a diverse tech stack (Go, Java, JavaScript, React GraphQL, Kubernetes, many public cloud web services, and more). As a senior backend engineer, you will help us build and extend next generation solutions such as a control plane for customers to manage their data pipelines at scale. New Relic is looking for engineers who are interested in building a brand-new observability experience. This high-impact engineering position is a phenomenal opportunity to own and build a set of next generation services and capabilities for the company. We are searching for a motivated engineer who is ready for a career-defining role in their next opportunity. We look forward to talking with you! What you'll do ● Design, Build, maintain, and scale back-end services and their support tools. ● Participate in architectural definitions with a high degr

javascriptjavareact
View job →
N
Newrelic
📍 San FranciscoFull-timeFrom $151K/yr
1mo ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity We are seeking an experienced and vision-driven Lead Enterprise Systems Engineer to join our engineering team. In this role, you will bridge the gap between business objectives, solution architecture, and hands-on execution. The ideal candidate remains actively involved in coding (roughly 70–80% of the time) while serving as the primary technical point of contact for project stakeholders. What you'll do Technical Vision & Solution Architecture Lead the architectural design, development, and deployment of resilient, scalable solutions in our Salesforce Platform for both Sales & CPQ. Translate business and product requirements into clear, technical roadmaps and system specifications. Establish engineering best practices, design patterns, coding standards, and testing strategies. Hands-On Execution & Quality Assurance Write clean, maintainable, and highly efficient APEX code alongside the Salesforce development team. Conduct thorough code reviews to ensure quality, security, and performance. Manage technical debt, proactively balancing speed of delivery with long-term system health. Team Leadership & Mentorship Provide technical guidance, direct support, and actionable feedback to Salesforce engineers. Mentor team members to foster technical growth and career advancement. Lead agile ceremonies (sprint planning, daily stand-ups, technical grooming, post-mortems). Cross-Functional Collaboration Partner closely with Technical Managers, Enterpri

ci/cdgitrest
View job →

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro

pythonsqlpostgresql
View job →
O
Okta
📍 IndiaFull-time$127K – $174K/yr · Jobiba est.
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
O
Okta
📍 IndiaFull-time$127K – $174K/yr · Jobiba est.
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
O
Okta
📍 IndiaFull-time$127K – $174K/yr · Jobiba est.
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams. Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning. Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity). Participate in t

pythonsqlpostgresql
View job →
O
Okta
📍 Bengaluru, IndiaFull-time$127K – $174K/yr · Jobiba est.
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. What You’ll Be Doing Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP. Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments. Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines). Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability. Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers. Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP. Champion SRE best practices — defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response. Lead complex technical projects from

pythonsqlpostgresql
View job →
P
Postman
📍 San FranciscoFull-time$3.1M – $3.3M/yr
1mo ago

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman’s AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally. What You’ll Do Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features Implement comprehensive observability and monitoring systems for real-time performance and fault detection Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence Drive continuo

aigorust
View job →
🔔

Get new lead devops engineer observability manager manager consultant jobs by email

Daily job updates · Unsubscribe anytime