Jobs in United States

Back End Td Reliability Lab Manager in United States

490 active opportunities · Updated October 2026

Explore current back end td reliability lab manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Datadog's Application Performance Monitoring (APM) provides deep visibility into the health, performance, and lifecycle of modern distributed applications, tracing requests from end-user devices (web and mobile) through to backend services. Our goal is to help customers detect root causes faster, optimize application performance, and improve resource efficiency at scale. As the Engineering Manager for APM Serverless, you will help define and deliver the end-to-end serverless APM experience, from auto-instrumentation through troubleshooting, and ensure that OpenTelemetry and Datadog-native customers alike have a frictionless and performant journey. You will also lead efforts to expand coverage of cloud-managed services across providers, ensuring customers can seamlessly trace and monitor critical services in all major and emerging cloud environments. We’re looking for an experienced engineering leader who thrives at the intersection of infrastructure and developer experience. You should care about well-designed APIs, observability-first thinking, and building systems that empower other developers. This is a high-leverage role that will influence how developers across the industry understand and instrument their serverless workloads. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Lead a polyglot team of 8-9 engineers and partner closely with Product and Engineering teams across Datadog to deliver industry-leading serverless capabilities that power consistent, scalable, and intuitive instrumentation across languages. Drive a domain that is technically rich: Lambda, Azure Functions, GCP, OTel billing, Rust, durable functions, distributed tracing across managed services. Engineers on this team work

AWSAzureGCPAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Full-Stack Engineer to join our Data Acquisition team to build and optimize the interfaces and tools that power our data infrastructure. Responsibilities: Develop and maintain full-stack applications that support data acquisition, including internal tools and dashboards. Collaborate closely with cross-functional teams, including Data Processing, Architecture, and Scaling, to ensure seamless data ingestion and workflow management. Design and implement APIs to facilitate data interactions between internal services and external data sources. Enhance user experience by developing intuitive web-based interfaces for managing and monitoring data pipelines. Optimize backend services for performance, scalability, and security in a distributed computing environment. Work with legal and compliance teams to ensure our data acquisition processes adhere to privacy regulations and best practices. Deploy and maintain infrastructure using Kubernetes and Infrastructure-as-Code (IaC) methodologies. Analyze system performance, conduct experiments, and improve data workflows to maximize efficiency. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in full-stack development. Proficiency in frontend frameworks (React, Vue, or similar) and backend technologies such as Python, Node.js, or Go. Strong expertise in RESTful APIs, GraphQL, and database design (SQL and NoSQL). Experience building data-intensive applications that handle large-scale datasets. Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker). Prior experience with web crawling and large-scale data processing is a

PythonReactNode.jsVue
P
📍 New York, NY, United States· Full-time· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a

PythonJavaC++Supply Chain
C
📍 Buffalo Grove, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Role Summary At CVS Health®, you’ll be working with a team of passionate colleagues who care deeply, innovate with purpose, hold themselves accountable and prioritize safety and quality in everything we do. Do you enjoy innovation while having fun doing it? If the answer yes, then this role might be for you! Join us and be part of something bigger, innovative and simplification in healthcare. We are seeking a highly experienced and innovative Principal (Director Level) Software Development Engineer to lead the application architecture, design, development, delivery of next-generation digital applications (including Reporting and financial solutions), and optimization of scalable, secure, and high-performance solutions leveraging AI across all major cloud platforms (AWS, Azure, and GCP). This role requires deep technical expertise, AI-enabled solutions, strategic thinking, scalable digital platforms, and enterprise integrations that power critical healthcare and pharmacy experiences and a passion for driving excellence in software engineering practices. This is a senior technical leadership role for a hands-on engineer who can operate across the full stack—from intuitive front-end applications to resilient backend services—while setting architectural direction, influencing engineering standards, and mentoring teams. The ideal candidate combines deep technical expertise, platform thin

TypeScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT engineering org builds and operates the systems that bring product improvements to users across backend services, web, mobile, and desktop platforms. The Developer Velocity team partners with product engineering, platform, infrastructure, reliability, engineering acceleration, and observability teams to make everyday development faster and releases safer, more predictable, and easier to operate. About the Role We are seeking a Technical Program Manager to improve developer velocity and deployment excellence across ChatGPT. You will lead durable improvements to local development, CI, testing, build systems, release trains, progressive rollout, and post-deployment validation. You will identify the highest-leverage sources of engineering friction, align teams around shared standards and metrics, and drive adoption of tooling and workflows that improve both speed and reliability. This role combines systems thinking, technical program leadership, and hands-on operating rigor across a broad engineering surface. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the cross-functional roadmap for improving local development, CI, testing, build workflows, and release infrastructure. Create a durable intake and prioritization mechanism for developer friction, using evidence to focus teams on the highest-impact improvements. Lead programs that improve deployment speed and safety, including pre-merge confidence, progressive rollout, release guardrails, rollback readiness, and post-deploy valida

AWSCI/CDRestAI
MT
📍 Richardson, TX, United States
✓ Quality checkedCompany trend +1266.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron’s High‑Bandwidth Memory (HBM) Digital Design organization is seeking a Digital Physical Design Engineer to drive backend implementation of high‑performance, low‑power digital logic used in advanced memory products. This role spans backend implementation from Netlist to GDSII, with a strong focus on timing closure, power integrity, and manufacturability in advanced process technologies. The successful candidate will work closely with RTL, DFT, CAD, and verification teams to deliver high‑quality physical designs, contribute to methodology improvements, and help push performance, power, and area (PPA) targets in complex, multi‑hierarchy designs. Key Responsibilities: Own and execute physical design implementation from synthesized netlist through GDSII, including floorplanning, placement, clock tree synthesis, routing, and physical signoff Utilize AI-Enabled tools in the Netlist to GDS flow Perform and drive timing closure across multiple modes, corners, and scenarios using industry‑standard STA tools Analyze and resolve congestion, timing, IR drop, EM, and power integrity issues Develop and refine floorplans, power grids, and clocking strategies for high‑performance designs Work with placement and timing closure of custom analog hard IP macros in a digital-on-top design flow Collaborate closely with RTL designers to influence partitioning, constraints, and micro‑architecture decisions early in the design cycle Implement and va

PythonAIRecruitment
C
📍 United States· Full-time
✓ Quality checkedCompany trend -94.7%

We’re looking for a Staff Software Engineer to help shape AI governance for developer tooling at Coder. This role sits on our AI Governance team, which builds and maintains two enterprise-grade components of Coder's AI governance stack. AI Gateway is a centralized LLM gateway that sits between coding agents and providers such as OpenAI or Anthropic, providing organizations with audit trails, token tracking, cost control, and centralized authentication. Agent Firewall wraps those agents with default-deny network policies, controlling which domains and methods they can reach inside workspaces. This team works across the full stack - from Go backend and React frontend to integrating with LLM provider APIs. Day to day, you'll be shipping features, hardening security boundaries, collaborating with enterprise customers on real-world policy needs, and contributing to Coder's open-source ecosystem. What you'll do here Design and build product features that push the standard for remote development in self-hosted environments Create and improve upon popular open source projects that integrate with VS Code, JetBrains, and other developer tools Champion best practices to both internal team members and external contributors Collaborate with Product and Design teams at Coder, as well as with partners like JetBrains, to execute key product integrations Document the design, implementation, and operations of systems for knowledge sharing within the team Work alongside Customer Success teams to support Coder’s enterprise user base Work with cutting-edge AI technologies to create seamless, painless developer experiences Rapidly iterate from prototype to implementation in a highly adaptive, reactive team environment What we're looking for 8+ years of full-stack experience writing code in a professional setting, with 1+ year(s) writing Go (ideally in current or most recent position) Proficiency in building distributed systems in Go Excellent verbal and written communication skills Excep

TypeScriptReactAWSGCP
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer - Multi-Agent Frameworks to join our AI Platform team. In this role, you will play a pivotal part in building a cutting-edge platform that empowers our users to create and deploy sophisticated intelligent agents, with a key focus on enabling collaborative and multi-agentic behaviors . This is a backend-focused role that requires deep expertise in AI, large language models (LLMs), and orchestration software. Key Responsibilities: Design, develop, and maintain a robust platform to enable users to create and manage AI agents and their interactions. Integrate and work with multiple LLMs, ensuring seamless orchestration and scalability for both individual and coordinated agent operations. Leverage orchestration frameworks like LangGraph and others to build complex workflows and pipelines that support diverse agent functionalities, including frameworks for multi-agent coordination . Develop and implement evaluation frameworks for testing AI agents in challenging and complex scenarios, focusing on individual performance and system-level dynamics. Stay at the forefront of AI advancements, incorporating the latest research and technologies into our platform to enhance agent capabilities and collaboration. Collaborate with cross-functional teams, including product managers, designers, and frontend engineers, to deliver a seamless user experience for building and deploying intelligent systems. Address challenging AI privacy scenarios, ensuring compliance with data protection regulations and best practices within agent-based applications. Contribute

AWSMachine LearningAI
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer - Multi-Agent Frameworks to join our AI Platform team. In this role, you will play a pivotal part in building a cutting-edge platform that empowers our users to create and deploy sophisticated intelligent agents, with a key focus on enabling collaborative and multi-agentic behaviors . This is a backend-focused role that requires deep expertise in AI, large language models (LLMs), and orchestration software. Key Responsibilities: Design, develop, and maintain a robust platform to enable users to create and manage AI agents and their interactions. Integrate and work with multiple LLMs, ensuring seamless orchestration and scalability for both individual and coordinated agent operations. Leverage orchestration frameworks like LangGraph and others to build complex workflows and pipelines that support diverse agent functionalities, including frameworks for multi-agent coordination . Develop and implement evaluation frameworks for testing AI agents in challenging and complex scenarios, focusing on individual performance and system-level dynamics. Stay at the forefront of AI advancements, incorporating the latest research and technologies into our platform to enhance agent capabilities and collaboration. Collaborate with cross-functional teams, including product managers, designers, and frontend engineers, to deliver a seamless user experience for building and deploying intelligent systems. Address challenging AI privacy scenarios, ensuring compliance with data protection regulations and best practices within agent-based applications. C

AWSMachine LearningAI
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Team Product Platform builds and owns the shared foundations the rest of Replit is built on, spanning the full stack so every other team can ship features safely and quickly: backend infrastructure, connectors, product primitives, and the frontend platform. Our work is high-leverage and horizontal: when our foundations are solid every other team moves faster, and the role gives you exposure across the whole of engineering. We are a small, collaborative team that values curiosity and clear thinking over pedigree, and we work in the open by bringing each other the problem rather than just the request. We care more about how you reason and build than the route you took to get here. About The Role As a Product Engineer , you can focus on frontend, backend, or full-stack work building the shared systems other teams depend on. The work is guided by a few simple questions: Are our shared systems fast, reliable, and cost-efficient as traffic grows? Are we making product development safe by default, consistent, and faster? Can a builder connect a third-party service once and have it work safely across every app they build? Are user-facing surfaces consistent and fast, with shared primitives teams can build on? Is our codebase easy to navigate, change, and extend, including for AI coding agents? What you’ll do Design reusable primitives and interfaces with clear contracts and documentation that other teams adopt Work directly with product teams to turn their friction into platform improvements Profile and instrument shared systems, then ship the improvements that move latency, cost, and reliability Harden systems against failure and abuse, and make safe defaults the path of least resistance Set technical direction in a

TypeScriptReactNode.jsRest
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Replit is a software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit helps more people build and ship software. About the Team Product Platform builds and owns the shared foundations the rest of Replit is built on: backend infrastructure, connectors, product primitives, and the frontend platform. When these foundations are solid, every other team moves faster. This role goes deep on the frontend platform: the architecture and platform layer behind every core product surface. The work is high-leverage and horizontal, and gives you exposure across the whole of engineering. We are a small, collaborative team that values curiosity and clear thinking over pedigree, and we care more about how you reason and build than the route you took to get here. If you like making other engineers faster, you will fit in here. About The Role As a Product Engineer focusing on the frontend platform , you will own the frontend architecture behind core product experiences: application frameworks, the API and data layer, testing infrastructure, and client performance. The goal is simple: product teams ship quickly and reliably on what you build. The team’s work is guided by a few simple questions: Is our core frontend architecture (frameworks, state, routing, SSR/CSR) sound, consistent, and easy to build on? Is our API and data layer reliable and ergonomic, with clear contracts, sensible error handling, and effective caching? Are user-facing surfaces fast and well-instrumented, with testing infrastructure that keeps them safe to change? Is the codebase easy to navigate, change, and extend, including for AI coding agents? You’ll partner closely with engineering, produc

TypeScriptReactNode.jsCI/CD
D
📍 Georgia, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $244K/yr

Quick readStrong listing-quality and freshness signals

The Language Tools team enables ~1,500 Datadog developers to build, test, and package millions of lines of Go, Python, Java, Rust, and TypeScript in our backend monorepo. Our success is measured by their productivity and satisfaction. They use the tools that we develop and support several times a day, in both development and CI environments. We use the Bazel open source build system as a foundation. The team is growing rapidly, both with Datadog and as we absorb other repositories into the monorepo. As a senior software engineer on the team, you will own projects from start to finish, both greenfield and brownfield. You will gain first-hand understanding of what Datadog developers need, and inform our roadmap. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Invent build, test and packaging tools that are simpler and more reliable to use. Push performance and cost efficiency at scale, raising cache hit rates and cutting CI times and compute spend across millions of targets. Treat CI like SREs treat prod, making sure our pipelines are green and fast. Prepare, run, and finish complex migrations. Contribute back to the Bazel ecosystem, upstreaming fixes and shaping features we depend on. Who You Are: An expert in Bazel and/or one of the languages listed above. A well-rounded engineer. You must broadly understand the various types of software projects that are built, tested, and packaged with our tools. Both careful and fearless. The changes we make impact the velocity of hundreds of engineers. They are risky but necessary. User-focused. We help Datadog engineers to use the tools that we develop, and continuously improve their usability, so they don’t need our help the next time. Ideally, you have ex

TypeScriptPythonJavaAI
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Senior Product Manager - Search Datadog’s Search team helps users – both human and agents – find answers to their questions. Search is a critical function in Datadog, and touches many product surfaces: the query editors in product homepages and dashboards, search bars for finding relevant assets, the global cmd+k navigation search, and the MCP tools that our Bits AI agent uses to respond to natural language prompts. Search is a full-stack team: owning user-facing search components, backend search ranking systems, and machine learning models to produce recommendations. The Search team is relatively new, and still growing. We have recently built out agentic search tools, and we are looking for a leader to help us expand to ambitious orchestration systems that deliver accurate results, and proactive recommendations across both keyword and semantic search. Beyond this, you’ll have room to influence how Datadog thinks about Search as a strategic surface. What you’ll do: Define and deliver how Datadog's AI agents discover the right context and tools to answer natural-language questions accurately and at scale Stay on top of industry trends in UI and agentic search experiences and capabilities Collaborate with Applied AI teams to integrate ranking, personalization, and recommendation models that scale across both human and agentic users Define and monitor KPIs for search quality, adoption, and downstream impact on user productivity; use them to drive data-informed decision making Engage directly with customers and internal product teams to deeply understand search journeys across query editors, global navigation, and natural-language agent interfaces Who you are: You have experience with search, ranking, or recommendation systems You are familiar with or very interested in MCP servers and differences between human and agentic UX You have a sharp eye for design and strong opinions on the micro-interactions — keyboard navigation, autocomplete behavior, loading

RestMachine LearningAIGo
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%

From $208.6K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . What we’re looking for: Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. In your role, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. You’ll grow as a person and leader in your field, all the while helping Pinners make their lives better in the positive corner of the internet. We are looking for a passionate, inquisitive, and well-rounded Sr. Staff Backend Engineer to join the Pinterest Assistant team. The team is building a visual-first, AI-powered companion that helps Pinners go from inspiration to action across shopping, search, and discovery. As the technical leader for our backend p

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist

AWSKubernetesRestAI
🔔

Get new back end td reliability lab manager jobs in United States by email

Daily job updates · Unsubscribe anytime