Jobs in United States

Senior Observability Engineer in United States

1,941 active opportunities · Updated October 2026

Explore current senior observability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Principal Engineer at Micron’s HBM New Product Validation team in Boise, ID you will be a senior technical guide responsible for validation strategy and readiness for future-generation High Bandwidth Memory (HBM) products. You will apply deep expertise in complex RAM and high-bandwidth memory architecture and building to establish architecture validation strategies. You will impact decisions that improve validation observability and debug efficiency . You will lead investigations of complex architecture, building , and silicon issues from pre-silicon simulation through post-silicon characterization. You will partner across Design Engineering, Design Validation, Product Engineering, Systems teams, and will develop and promote AI-enabled validation methodologies that improve efficiency, quality, and scalability. Responsibilities: Define architecture-aware validation strategies and coverage direction for future-generation HBM products. Drive validation readiness for next-generation products by preparing strategies, test environments, and execution plans before first silicon, reducing risk, accelerating issue resolution, and preventing late-stage surprises. Evaluate new DRAM/HBM architectural features and define validation requirements and testability needs before DBR. <span style="color:#

PythonGitLinuxAI
M
📍 O Fallon, Missouri, United States
✓ High-confidence listingCompany trend +212.5%
Quick readStrong listing-quality and freshness signals

Our Purpose Mastercard powers economies and empowers people in 200&#43; countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Principal Software Engineer Principal Software Engineer Position Overview The Principal Software Engineer is a senior individual contributor within Architecture & Technology (A&T), responsible for driving enterprise engineering strategy, architecture standards, and technology excellence across Mastercard. This role provides deep technical leadership across distributed systems, cloud-native platforms, resiliency, observability, and software engineering practices while influencing technology direction across multiple teams and domains. The Principal Engineer partners with senior engineering leaders, architects, and platform organizations to define architectural standards, establish reusable patterns, and guide critical technology decisions. Through technical expertise, thought leadership, and cross-functional influence, this role helps teams build secure, scalable, reliable, and operationally excellent solutions that align with Mastercard's long-term engineering strategy. Role • Provide technical leadership and architectural guidance across multiple engineering organizations, driving consistent adoption of engineering standards, best practices, and enterprise technology patterns. • Partner with platform CTOs, architects, and engineering leaders to evaluate technology investments, transformation initi

AWSAzureGCPAI
D
📍 Massachusetts, California, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $211K/yr

Quick readStrong listing-quality and freshness signals

We are Datadog's in-house product experts. The technical solutions team enables Datadog's worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. The Product Solution Architecture team works closely with Datadog customers, Datadog’s Product management, engineering and Tech Solutions teams to help architect and implement the Datadog Solution. The team has grown to cover more than 12 Product Categories and needs a highly motivated, results-oriented individual to help lead the team. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Manage a team of expert individual contributors, working with Engineering and Product Management on their Product, who help create best practices for adoption of Datadog at scale. Assist in recruiting and hiring of top talent for the team. Mentor/coach new hires during on-boarding to ensure proper ramping of skills and capabilities Ensure that your team is aligned to top company initiatives and enabled to work with various Datadog internal and external stakeholders Deliver performance reviews along with collaborating on and executing individual development plans Be a Trusted Advisor to Product Management by providing high quality feedback based on your field experience working closely with our Sales, Support, Partners and Customers Offer guidance on architecture choices, data collection, and best practices to large Enterprise customers as they adopt these Datadog services across the organization Participate in Executive Brief conversations with high profile customers and prospects about overall architecture, Observability Best Practices and industry trends Who You Are: 10+ years of experience solving complex problems for customers and a strong kno

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%
Quick readStrong listing-quality and freshness signals

Datadog's integrations are the connective tissue between our platform and the technologies our customers run in the real world. As a Sr. PM on the Agent Integrations team, you will own the vision, prioritization, and execution for 100+ integrations that run directly inside the Datadog Agent from foundational infrastructure (MySQL, Kafka, Kubernetes) to the rapidly growing landscape of self-hosted AI and on-premise enterprise technologies. This is a high-impact, breadth-first role at the intersection of infrastructure observability and the frontier of AI-native workloads. At Datadog, we place value in our office culture; the relationships it builds, the creativity it brings, and the collaboration of being together. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Own the Agent Integrations roadmap. Determine which new integrations to build and which existing ones to improve, balancing customer demand, business impact, and engineering capacity across a catalog of 100+ technologies. Drive the expanding AI integration surface. Lead product strategy for self-hosted AI workloads, including LLM inference frameworks (e.g., Hugging Face TGI, BentoML), AI agents, MCP servers, and model orchestration tools, so Datadog customers can monitor every layer of their AI stack. Expand on-prem and hybrid coverage. Prioritize and execute new integrations for on-prem technologies including storage systems, HPC schedulers, network devices, and legacy enterprise platforms where customers run critical workloads. Build observability for ERP systems. Define and drive Datadog's strategy for monitoring enterprise ERP platforms (SAP, Oracle EBS/Fusion, Microsoft Dynamics) covering performance, job execution health, and integration layer telemetry so enterprise customers can observe their ERP stack alongside the rest of their infrastructure. Analyze adoption and customer feedback at scale. Use data from multiple sources to

SQLPostgreSQLMySQLMongoDB
O
📍 Atlanta, Georgia, United States· Full-time
✓ High-confidence listing

From $116.5K/yr

Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this role, you will part of the R&D Team that works on mission-critical applications. Your Mission Engage and partner with various Engineering, Operations, and Product teams to design, deliver, and maintain a highly available and performant application platform. Build and implement application observability and platform monitoring tools to continuously improve the customer experience Eliminate toil by automating processes, tuning alerts, and improving code where it is most needed Frequently evaluate new ideas and trends to identify potentially useful tools and techniques Collaborate with different functional groups to identify gaps, prioritize, and resolve issues Defining, implementing, and maintaining SLIs and SLOs aligned with customer experience. Design and instrument SLIs such as latency, error rates, and availability across critical services Manage and enforce error budgets to balance system reliability with product feature v

PythonJavaSQLAWS
O
📍 Atlanta, Georgia, United States· Full-time
✓ High-confidence listing

From $116.5K/yr

Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this role, you will part of the R&D Team that works on mission-critical applications. Your Mission Engage and partner with various Engineering, Operations, and Product teams to design, deliver, and maintain a highly available and performant application platform. Build and implement application observability and platform monitoring tools to continuously improve the customer experience Eliminate toil by automating processes, tuning alerts, and improving code where it is most needed Frequently evaluate new ideas and trends to identify potentially useful tools and techniques Collaborate with different functional groups to identify gaps, prioritize, and resolve issues Defining, implementing, and maintaining SLIs and SLOs aligned with customer experience. Design and instrument SLIs such as latency, error rates, and availability across critical services Manage and enforce error budgets to balance system reliability with product feature v

PythonJavaSQLAWS
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

PythonCI/CDGitRest
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $180K/yr

Quick readStrong listing-quality and freshness signals

As a Director, Enterprise Sales Engineering, you will lead a team of frontline leaders, driving account strategy, deal execution, operational excellence, and team development across your region. This is a strategic leadership role with regional impact within North America. You will act as a force multiplier by owning team and region performance, shaping execution standards, and partnering closely with Sales leadership to influence go to market strategy, improve outcomes, and scale impact beyond individual deals. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own in-field execution at a regional level by taking an executive presence role in strategic and high-impact opportunities and providing senior-level stakeholder alignment and technical credibility alongside your team. Develop and operationalize a repeatable formula and methodology that can scale across a large customer base, enabling Datadog to accelerate enterprise growth by illustrating to customers how Datadog’s observability and security solutions quickly, simply, and efficiently deliver measurable business outcomes especially in a world increasingly dominated by Agentic AI. Partner confidently with senior Sales leadership on regional strategy, multi-quarter planning, annual capacity planning, and forecast alignment operating 2-3 quarters ahead to anticipate resourcing needs and business priorities. Oversee consistent, high-quality delivery across technical sales cycles in the region, including evaluations, customer engagements, and deal progression. Lead, coach, and develop a team of Sales Engineering leaders, driving performance, engagement, and long-term career growth through regular feedback and structured development. Hire and onboard top talent across ICs and manager-lev

AWSAzureGCPKubernetes
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $128K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor

PythonKubernetesLinuxAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for an exceptionally experienced, hands-on full-stack engineer to define and build the next generation of AI-powered enterprise workflows. You will take on the hardest and most ambiguous problems in the Technology vertical: translating real customer needs into product direction, designing the systems behind the experience, and personally writing and shipping production-quality code across the stack. You will own the technical direction and end-to-end delivery of products spanning ChatGPT Work surfaces, backend services, plugins, connectors, enterprise data, permissions, and evaluations. You will make foundational architecture and product tradeoffs; establish patterns other engineers can build on; and hold these experiences to a high bar for reliability, security, observability, and customer value. This is an individual-contributor role for an engineer who leads through technical judgment, direct execution, and influence—not people management. You should be equally comfortable working directly with customers, setting direction with senior cro

TypeScriptPythonReactNode.js
O
📍 Mountain View, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Monetization Data Systems team builds the trusted data and product systems that power how the company develops, measures, and improves monetization products. We bring together product usage, pricing, billing, ads, payments, and financial data to help Product, Engineering, Finance, and GTM teams make better decisions and deliver reliable customer experiences. We work at the intersection of data engineering, product engineering, platform engineering, Finance, and GTM. Our goal is to turn complex monetization and financial data into accurate, explainable, and timely data products while building systems that scale with the growth and complexity of the business. About the Role We are looking for a Senior Software Engineer to design and build the next generation of our monetization data platform. You will own high-impact platform systems end to end, from architecture and implementation through testing, deployment, observability, and ongoing operation. This is a hands-on role for an engineer who enjoys solving ambiguous customer and business problems, designing durable systems, and partnering closely with Product, Data, Finance, Accounting, and GTM. You will help define technical direction, raise the engineering bar, and turn monetization opportunities into reliable, scalable product experiences and platform capabilities. In this role, you will: Design, improve, and operate reliable, scalable backend services that power pricing, billing, ads, payments, entitlements, and other monetization platform capabilities. Own the architecture and implementation of critical workflows relevant to monetization data, data contracts, and integrations across product and business systems. Establish strong guarantees for correctness, availability, security, performance, reconciliation, and auditability across business-critical systems. Build reusable platform capabilities and developer tools that enable product teams to launch, measure, and iterate on monetization products

PythonJavaAWSRest
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. As a Senior Product Designer, you'll own the experience for some of the most high-stakes, technically complex workflows in enterprise softwa

C
📍 Woonsocket, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. At CVS Health, Site Reliability Engineering (SRE) is fundamental to delivering the reliable, secure, and scalable technology experiences that support millions of patients, customers, pharmacists, and healthcare professionals every day. Our SRE organization drives operational excellence across critical healthcare and retail platforms through innovation, automation, observability, and engineering best practices. The Executive Director, Site Reliability Engineering serves as the strategic leader responsible for the reliability, resilience, and performance of CVS Health's retail and pharmacy technology ecosystem. This executive will define and execute a comprehensive reliability strategy, oversee large global engineering teams, and establish a long-term vision for observability, automation, and operational excellence across thousands of store locations. Working closely with senior business and technology leaders, the Executive Director will champion modern SRE practices, accelerate incident response capabilities, and deliver real-time operational visibility that enables proactive issue prevention and exceptional customer and patient experiences. Key Responsibilities Strategic Leadership & Vision Define and lead the enterprise-wide Site Reliability Engineering strategy supporting CVS Health's retail and pharmacy operations. Align reliability and operational

AWSAzureGCPKubernetes
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $154K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

PythonKubernetesAISwift
B
📍 Raleigh, North Carolina, United States
✓ High-confidence listingCompany trend +350%
Quick readStrong listing-quality and freshness signals

This is where your work makes a difference. At Baxter, we believe every person—regardless of who they are or where they are from—deserves a chance to live a healthy life. It was our founding belief in 1931 and continues to be our guiding principle. We are redefining healthcare delivery to make a greater impact today, tomorrow, and beyond. Our Baxter colleagues are united by our Mission to Save and Sustain Lives. Together, our community is driven by a culture of courage, trust, and collaboration. Every individual is empowered to take ownership and make a meaningful impact. We strive for efficient and effective operations, and we hold each other accountable for delivering exceptional results. Here, you will find more than just a job—you will find purpose and pride. Your Role at Baxter Provides enterprise-level technical leadership for cloud shared services and related connected-care ecosystems. Collaborates with engineering, product, and business leaders to shape long-term architectural strategy, establish technical standards, and guide the evolution of secure, cloud-native services that support multiple products, regions, and business domains. Serves as a senior technical leader for distributed systems, GraphQL and API architecture, multi-region cloud strategy, service interoperability, scalability, security, resiliency, observability, SOC 2 readiness, and operational excellence. Partners closely with executive leadership, product management, cybersecurity, quality, regulatory, operations, and engineering teams to align technology investments, architectural decisions, and platform capabilities with business objectives and sustained growth. <span style="color:

Node.jsAzureKubernetesGraphql
🔔

Get new senior observability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime