Jobs in United States

Back End Td Reliability Lab Manager in United States

490 active opportunities · Updated October 2026

Explore current back end td reliability lab manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Fullstack Software Engineer, you will design and build the systems and experiences that power how millions of people connect to their finances. You will work across the stack, building scalable backend services and APIs while also crafting intuitive, high-quality frontend experiences that bring those systems to life. This role is ideal for engineers who enjoy switching between backend problem-solving and frontend user experience work, and who are excited to grow their impact across both. You will collaborate closely with product managers, designers, and other engineers to ship products that are reliable, secure, and delightful to use. At Plaid, engineers take ownership early, contribute to architectural decisions, and see their work reach millions of users. Responsibilities: Build across the stack. Design, develop, and maintain scalable backend services and APIs, as well as intuitive, high-quality frontend applications that bring those systems to life. Collaborate cross-functionally. Partner closely with product managers and designers to define requirements and de

JavaScriptJavaSQLMySQL
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e

AWSGitMachine LearningAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e

AWSGitMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for a full-stack product engineer who can own ambiguous enterprise workflows end to end: understand a customer problem, shape the product, build across frontend and backend, work through platform dependencies, instrument quality, and learn quickly with design partners. You will build across ChatGPT Work surfaces, services, plugins, connectors, and data or permission boundaries when the experience requires it. You will make quality and rollout observable through evaluations, product and operational signals, and clear fallback or rollback paths. This is a product-engineering role for someone who can move between user problems and system details without losing ownership of either. The strongest candidates will be able to turn specific customer evidence into a generalizable product, explain scope and architecture tradeoffs, and carry a feature from an early prototype through a bounded production rollout. In this role, you will: Build and ship role-specific workflows across ChatGPT Work surfaces, services, plugins, and connectors. Turn customer a

TypeScriptPythonReactNode.js
S
📍 Nyc, United States· Full-time· Remote
✓ Quality checkedCompany trend -100%

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We're looking for a Developer Relations Engineer based in New York, NY to join our team and help more developers discover, learn, and build with Supabase. You'll create high-impact content, build real-world projects, and represent Supabase across communities and events. If you're equally energized by writing code and teaching others, this is the role for you. Why This Role Matters Supabase is growing fast with 350,000+ developers , 1,000+ OSS contributors , and a thriving open source ecosystem. Our users are builders, startup founders, weekend hackers, and engineers scaling to millions of users. DevRel is how we meet them where they are: through content, community, and code. We're building a community of communities that brings together developers from many backgrounds, including first-time open source contributors. As a DevRel Engineer, you'll be a bridge between Supabase and the broader developer ecosystem, helping people get started, go deep, and feel connected. What You'll Do Make content — Publish compelling technical content, especially video, to help developers learn Supabase quickly Build demos and tutorials — Ship real-world apps using Supabase and tools like Next.js, React, and Stripe. Write guides that others can follow and remix Represent Supabase — Speak at meetups, livestream builds, and engage with the developer ecosystem. You'll be a visible and trusted voice of the platform Support and grow the community — Celebrate contributors, answer questions, highlight cool projects, and bring developer feedback to the team Collaborate across the company — Work with engineering, product, and growth to amplify launches, prioritize content, and reduce friction for new users Yo

JavaScriptTypeScriptJavaReact
S
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase is looking for an Event Programs Manager, Developer Community & Ecosystem to help grow and scale our developer community initiatives and industry event programs. You’ll work closely with the Head of Global Events and Field Marketing to build programs that connect Supabase with developers and the broader ecosystem. This includes initiatives such as meetups, workshops, hackathons, and community activations that bring Supabase into the spaces where developers learn, build, and collaborate. As an early member of the events team, you’ll help shape how Supabase shows up in person around the world while building the playbooks and frameworks that allow these programs to scale globally. This role sits within Field Marketing and works cross-functionally across Marketing, DevRel, Partnerships, Sales and Design to support programs that deepen developer engagement, strengthen ecosystem relationships. You will also support programming around Supabase Select, our flagship user conference, including partner and community programming connected to the event. We’re a lean, remote-first team, so this role requires someone who is comfortable operating independently, experimenting with new ideas, and helping build the systems that will power Supabase’s events program as it grows. What You’ll Do Own and Scale Community Programs Scale Supabase developer meetups, workshops, and hackathons globally Partner with DevRel to align community events with product launches and developer education initiatives Coordinate community meetups around major industry events and Launch Week Support community programming connected to Supabase Select Support the SupaSquad advocate network through community activ

S
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase's Partnerships function is scaling fast - spanning Technology Partners, Solution Partners, Startups, and Cloud Partnerships. We're looking for a Partner Operations & Systems Lead to own the operational and technical backbone that lets this team move quickly and make decisions with good data. This is a senior, hands-on role. You'll design the systems, processes, and reporting that the whole partnerships org runs on - and, increasingly, you'll build the internal tools yourself. Supabase is investing heavily in AI-assisted development to move faster than a traditional ops build cycle allows, and this role is expected to be a leading example of that inside Partnerships. You'll report directly to the Head of Partnerships and sit alongside our regional and functional leads as a peer, with the mandate to build and enforce operational rigor across all teams. Beyond the infrastructure, we want someone who helps shape where Partnerships places its bets - not just the system that reports on them. What You'll Own Systems & Data Infrastructure Own the partner tech stack end-to-end - CRM partner objects, attribution tooling, and partner-facing portals. Design and maintain partner attribution models and the dashboards leadership uses to evaluate performance. Ensure data hygiene and consistency across partner records, deals, and touchpoints spanning all sub-functions. Process & Program Management Design and run partner onboarding, tiering, and certification programs. Own the RFC / DRI decision-making framework for the partnerships team, including how it's used and refined over time. Run deal-registration and "quarterback" account-ownership processes that keep internal teams

S
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We’re looking for a Developer Relations Engineer based in San Francisco to join our team and to help more developers discover, learn, and build with Supabase. You’ll create high-impact content, build real-world projects, and represent Supabase across communities and events. If you’re equally energized by writing code and teaching others, this is the role for you. Why this role matters Supabase is growing fast with 350,000+ developers , 1,000+ OSS contributors , and a thriving open source ecosystem. Our users are builders, startup founders, weekend hackers, and engineers scaling to millions of users. DevRel is how we meet them where they are: through content, community, and code. We’re building a community of communities that brings together developers from many backgrounds, including first-time open source contributors. As a DevRel Engineer, you’ll be a bridge between Supabase and the broader developer ecosystem, helping people get started, go deep, and feel connected. What you'll do Make content Publish compelling technical content, especially video, to help developers learn Supabase quickly Build demos and tutorials Ship real-world apps using Supabase and tools like Next.js, React, and Stripe. Write guides that others can follow and remix Represent Supabase Speak at meetups, livestream builds, and engage with the developer ecosystem. You’ll be a visible and trusted voice of the platform Support and grow the community Celebrate contributors, answer questions, highlight cool projects, and bring developer feedback to the team Collaborate across the company Work with engineering, product, and growth to amplify launches, prioritize content, and reduce friction for new users You migh

JavaScriptTypeScriptJavaReact
L
📍 San Diego, California, Canada
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Ignite your curiosity. Solve the unsolvable. At Leidos, we do more than write code—we decode the unknown. Our San Diego-based research and engineering team takes on some of the nation’s toughest defense challenges using advanced signal processing, ocean remote sensing, and high-performance computing. We’re seeking a Software Engineer / Computer Scientist who enjoys solving complex problems and pushing the limits of performance. In this role, you’ll work alongside a multidisciplinary team of scientists and engineers with expertise in hydrodynamics, physics, acoustics, and signal processing to build impactful software that turns massive, complex data sets into meaningful insight. If you are motivated by innovation, energized by collaboration, and excited to see your work support real-world missions, this could be the right opportunity for you. What You’ll Do Collaborate with scientists and engineers to design, develop, and optimize advanced algorithms for next-generation radar, optical, and infrared sensor systems. Build scalable, high-performance backend systems for scientific computing in distributed environments. Integrate, refactor, and improve scientific codebases to increase efficiency and scalability. Translate and optimize existing code for GPU/CUDA acceleration and parallel or distributed execution. Test, document, maintain, and enhance complex software in Linux/Unix environments. Contribute in a collaborative environment that values technical excellence, creativity, and continuous growth. Required Qualifications Bachelor’s degree in Computer Science, Applied Mathematics, Physics, or a related field with 4+ years of backend software development experience, or a Master’s degree with 2+ years of experience. Equivalent experience may be considered in place of a degree. U.S. citizenship and the ability to obtain a Top Secret clearance; active Top Secret clea

GitLinuxAI
D
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $100K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data structures,

KubernetesRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Customer education helps customers and partners build the practical skills and confidence to use AI and OpenAI products safely and effectively. The team focuses on role- and skill-based learning paths, practical content, and product experiences that accelerate learning in the workplace. It brings together learning and enablement expertise, field insight, product signals, and measurement to improve learner and business outcomes. Together, these experiences will help enterprise users build practical AI skills, apply them with confidence in their work, and demonstrate what they can do. Employers will gain a clearer view of workforce skills and progress, helping them recognize capability, focus development where it matters most, and build confidence in workforce readiness. About the Role We’re looking for a full-stack engineer to define and build a new class of learning experiences. This is an early-stage product area where technical judgment, product sense, and learner empathy are critical. You will be setting a technical vision for how people use AI to learn how to use AI, safely and beneficially. This is a hands-on, 0-1 product engineering role with broad technical and product ownership. You’ll set direction, make foundational decisions, and ship the first versions of experiences that can grow into the default way people learn at work. You will drive full-stack product experiences end to end, from prototype through launch, instrumentation, iteration, and production hardening. The work spans interaction design, frontend implementation, backend APIs and services, learner state, content and runtime integration, telemetry, evaluation, reliability, safety, accessibility, and launch readiness. You’ll work closely with our education, GTM, and engineering teams to translate how people learn into products people want to use. bring role- and skill-based learning paths into the product, designing coaching, feedback, and adaptive support which responds to each lea

TypeScriptReactAWSRest
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay curre

TypeScriptPythonAWSAzure
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer – AI Platform to join our ClickUp Engineering team. In this role, you will play a critical part in both building the core AI platform and directly applying large language models (LLMs) to deliver intelligent features across ClickUp. You will focus on backend systems that enable scalable, reliable, and secure AI-powered capabilities, while also working hands-on with LLMs to solve real user problems and drive product innovation. Key Responsibilities: Architect, design, and implement scalable AI platform services that support the deployment, orchestration, and lifecycle management of LLMs and other AI models. Apply LLMs and other AI technologies directly to build and enhance ClickUp’s intelligent features, working closely with product and engineering teams to deliver impactful solutions. Build and maintain robust APIs and backend systems that enable seamless integration of AI-powered features into ClickUp’s core platform. Develop infrastructure for model serving, monitoring, logging, and automated evaluation to ensure high reliability and performance of AI services in production. Integrate with multiple LLM providers (e.g., OpenAI, Anthropic, Google) and manage model selection, routing, and fallback strategies for optimal performance and cost. Drive the adoption of best practices in AI privacy, security, and compliance, including data anonymization, secure data handling, and regulatory adherence. Optimize platform performance, scalability, and cost-efficiency, leveraging cloud-native technologies and distributed systems. Stay current with ad

TypeScriptPythonAWSAzure
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Job Summary We are looking for an experienced Growth Infrastructure Engineer to build and maintain the technical backbone that enables scalable growth experiments, high-performance data pipelines, and automated systems that drive user acquisition, engagement, and product iteration. This role sits at the intersection of growth, product, and infrastructure — combining deep technical engineering with experimentation and data-driven optimization. You will collaborate with product, data science, and backend teams to ensure that growth initiatives run smoothly and scale efficiently across systems. Key Responsibilities Growth Infrastructure & Systems Design, implement, and maintain scalable infrastructure that supports growth and experimentation needs. Build and optimize analytics pipelines to capture key product and growth metrics (acquisition, activation, retention, etc.). Develop automated workflows for user onboarding, campaign delivery, and performance tracking. Experimentation & Optimization Support A/B testing frameworks and integrate them into production systems. Enable reliable data collection and evaluation for growth experiments. Automate deployment and rollout of growth feature flags and tests. Cross-Functional Collaboration Partner with Growth Product Managers, Data Engineers, and Analysts to define technical requirements for growth initiatives. Translate business goals into technical specifications and system designs. Provide guidance on performance, reliability, and scalability trade-offs. Monitoring & Reliability Implement monitoring and alerting for growth infrastructure services. Troubleshoot production issues and optimize for uptime and performance. Ensure data quality and consistency for report

JavaScriptPythonJavaAWS
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $100K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data struc

KubernetesGitRestAI
🔔

Get new back end td reliability lab manager jobs in United States by email

Daily job updates · Unsubscribe anytime