Jobiba hiring network

Back End Td Reliability Lab Manager Jobs

1,726 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current back end td reliability lab manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.

PE
Private Employer
📍 Bengaluru• Full-time• Hybrid
1mo ago

About Bazaarvoice At Bazaarvoice, we create smart shopping experiences. Through our expansive global network, product-passionate community & enterprise technology, we connect thousands of brands and retailers with billions of consumers. Our solutions enable brands to connect with consumers and collect valuable user-generated content, at an unprecedented scale. This content achieves global reach by leveraging our extensive and ever-expanding retail, social & search syndication network. And we make it easy for brands & retailers to gain valuable business insights from real-time consumer feedback with intuitive tools and dashboards. The result is smarter shopping: loyal customers, increased sales, and improved products. The problem we are trying to solve : Brands and retailers struggle to make real connections with consumers. It's a challenge to deliver trustworthy and inspiring content in the moments that matter most during the discovery and purchase cycle. The result? Time and money spent on content that doesn't attract new consumers, convert them, or earn their long-term loyalty. Our brand promise : closing the gap between brands and consumers. Founded in 2005, Bazaarvoice is headquartered in Austin, Texas with offices in North America, Europe, Asia and Australia. It’s official: Bazaarvoice is a Great Place to Work in the US , Australia, India, Lithuania, France, Germany and the UK! What you’ll be doing Lead, hire and grow a high-calibre team of frontend & backend engineers, and their line managers. Define and implement a roadmap based on critical business need, that delivers the valuable features, scale, and reliability our clients need, as well making systems more resilient and scalable. Drive business-significant and complex initiatives by collaborating across geographically distributed teams and partners Coach and mentor engineers globally in support of their growth and adherence to best practices. Who you are – Requirements for success in this rol

awsci/cdrest
View job →
F
1mo ago

About Us What if your work could drive change in a globally established industry, shaping processes that touch every corner of the world? At Forto, we are at the forefront of change, harnessing the power of AI to revolutionise logistics. We want to reinvent digital supply chains to be transparent, frictionless and sustainable. From day one, our mission has been to simplify global trade – creating a seamless and efficient logistics process. Your role & Mission The mission of the Flash team at Forto is to fundamentally redesign operations processes by leveraging automation and intelligent decision-making to handle shipments more efficiently. We aim to scale CoPilot so that it becomes the primary system used by operations managers, enabling them to be more effective and efficient in their daily work. As a Senior Software Engineer in the Flash team, you will help build AI-driven solutions and the CoPilot that powers our logistics operations. You will maintain and evolve a sophisticated event-driven, distributed architecture designed to answer one key question: How do we improve shipment handling and bring efficiency to our operations teams at scale? From quotation and rate management, to shipment execution and schedule optimization (capacity utilization and GP optimization), to carrier integrations and automated data extraction from unstructured sources (emails, PDFs, spreadsheets), this role focuses on building reliable, data-heavy systems that directly impact revenue and operational performance. What you will do Design, build, and evolve scalable and resilient backend systems Contribute to an event-driven, distributed architecture Work on AI-adjacent systems (integration, orchestration, data-heavy workflows) Own services end-to-end (design, implementation, documentation, operation) Collaborate closely with product managers, operations, and other engineers Communicate clearly with stakeholders and explain technical trade-offs Contribute to technical standards and b

typescriptnode.jsmongodb
View job →
R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. What You’ll Do: As a Senior Backend Engineer on the Avatar Marketplace team, you will be building an international marketplace for millions of virtual items that users acquire to customize their avatars and identities on the Roblox platform. You will design, implement, and operate the frameworks and services you and the team build that help control the supply and demand of the Roblox economy. The team is responsible for building systems that will help content creators become more successful sellers. With over 90 million daily active users (and growing), we are looking for an experienced engineer who is passionate about designing and building scalable systems for both sides of the marketplace. The Marketplace Foundation team's focus is to provide reliable and scalable solutions for critical systems supporting both Avatar Marketplace creators and users. We aim to invent new and evolve existing core concepts to support new and future use cases in a holistic manner, such as Bundles, Limiteds, and Price Optimization. We help the creators to improve their earning while scaling Roblox's economy. The team is also responsible for ensuring 10x scalability headroom on critical Marketplace system

pythonjavamongodb
View job →
T
Taskrabbit
📍 San Francisco• Full-time• $106K – $142K/yr
1mo ago

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role operates on a hybrid schedule requiring two days of in-office collaboration per week. The position must be based in the San Francisco Bay Area. About the Role We're hiring a Software Engineer II within our Fulfillment organization — the backend systems that get the right job to the right Tasker and see it through to completion. You'll join Fulfillment Lifecycle, the team that decides how jobs are matched to Taskers for our partner and marketplace business, increasingly using unstructured data and experimentation to make matching smarter and fulfillment more reliable. The team is part of a company-wide platform modernization effort, breaking a legacy monolith into well-bounded, API-first services. We're hiring for a strong backend engineer who thrives on complex, data-intensive problems, is comfortable with ambiguity, and takes pride in well-tested, observable, production-ready code. What You'll Work On B

typescriptnode.jsaws
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Senior Software Engineer, Simple Trade Experience We're hiring a Senior Software Engineer to join the Simple Trade Experience team within the Consumer & Business group. This team powers the critical buy, sell, and convert journeys on the Simple interface of the Coinbase app and website, supporting multiple trading types and assets for millions of users worldwide. You'll own highly performant, available, and consistent backend and frontend systems that directly enable customers to trade with confidence at web-scale. What you'll do: Own the design and delivery of highly performant, available, and consistent systems powering trading for millions of users worldwide Architect robust, extensible systems that support multiple trading types and assets (Spot, Limit, Recurring, and more) while maintaining reliability at web-scale Partner closely with product and design to define and ship a best-in-class trading experience across web and mobile Drive production service evolution, scaling and maintaining critical trading infrastructure as traffic and product scope grow Strengthen engineering quality across the team through rigorous code reviews, mentorship, and knowledge sharing Required Skills and Experience: 5+ years of experience in software engineering with demonstrated success building and scaling production services handling web-scale traffic Proven track record design

REMOTEjavamongodbredis
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. Apply and scale optimization techniques across a wide range of ML models, particularly large language models. Collaborate with a diverse team to design and implement innovative solutions. Own projects from idea to production. REQUIREMENTS Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. Experience with one

pythondockerkubernetes
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for a seasoned Frontend Engineer to craft performant and delightful user experiences across Baseten’s core platform. You’ll own critical parts of our web application stack and collaborate cross-functionally with product, design, and backend teams to launch impactful features that help users deploy and manage AI systems at scale. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Core Product team: Rolling Deployments Model APIs for frontier models Model training built for production inference RESPONSIBILITIES Design, implement, and maintain responsive, accessible, and user-friendly frontend interfaces using React and TypeScript Collaborate closely with product designers to turn complex ideas into elegant, intuitive UIs Optimize application performance and reliability, with a focus on rendering speed and responsiveness Drive major frontend initiatives, including partnering with backend teams to define APIs and test and refine end-to-end flows Establish best practices, and mentor other engineers on frontend technologies Build reusable component libraries and frontend infrastructure that accelerate product development Partner with backend and platform teams to define and refine APIs and end-to-end flows REQUIREMENTS 5+ years of experience building production-grade web applications Deep expertise in React, TypeScript, and modern web development tooling Track record of bu

typescriptreactmachine learning
View job →
S
1mo ago

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. We’re looking for a Principal Engineer to join the ML Platform team at Synthesia. Our team builds and operates the systems that allow researchers and product teams to train, serve, and deploy generative models reliably and efficiently . This includes research infrastructure, production serving systems, internal tooling, and the platform interfaces that connect them. A growing part of our mission is making these systems more automation-friendly and agent-oriented , so that workflows can increasingly be operated through reliable tooling rather than manual effort. We’re looking for a strong generalist with a systems mindset: someone who is comfortable working across infrastructure, backend systems, and tooling, and who has seen ML systems in practice. this is not a pure ML Engineer role. We’re especially interested in people who think deeply about reliability, scalability, performance, and resource efficiency in complex production environments. This is a hands-on IC role with significant ownership. You’ll help shape how our ML platform evolves as we scale the number of models, workloads, tools and teams relying on it. What you’ll do Design and improve the platform systems that support model training, evaluation, an

pythonkubernetesgit
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: Optimizing performance of large-scale workloads on Ray Stability and stress testing infrastructure Improving fault tolerance (HA) As part of this role, you will: Leading cross-team projects while mentoring junior team members Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements

restmachine learningai
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Ray aims to provide a universal API for building distributed applications (e.g. a machine learning pipeline of feature engineering, model training, and evaluation). Data is usually a core element connecting these different stages, and therefore plays a critical role in Ray’s usability, performance, and stability. We are looking for strong engineers to build, optimize, and scale Ray’s Datasets library and data processing capabilities in general. About the Ray Data team: The Ray Data team currently develops and maintains the Ray Datasets library, which is already powering critical production use cases (e.g. large scale data compaction at Amazon , and ML pipeline at Alibaba ). Ray Datasets is a Python library built on top of Apache Arrow and Ray Core (Ray’s C++ backend), and the Ray Data team interacts closely with Ray Core components including the scheduler and the memory & I/O subsystems. The Ray Data team also works closely with Ray’s ML libraries including Train, RLlib, and Serve. A snapshot of projects you will work on: - Performance of Ray Datasets at large scale (leveraging Arrow primitives, optimizing Ray object manager, etc.) - Integration with ML training and data sources - Stability an

pythonmachine learningai
View job →
A
Anyscale
📍 Remote• Full-time
1mo ago

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role Ray aims to provide a universal API for building distributed applications. To achieve this goal requires a distributed system with high levels of performance and reliability. We're looking for engineers with systems software experience that are interested in contributing to the Ray backend. About the Ray Core Team The Ray Core team develops and maintains the Ray C++ backend (e.g., distributed scheduler, language runtime integration, I/O and memory subsystems). We are responsible for the reliability, scalability, and performance of Ray as well as ensuring that Ray provides the right feature set to support higher level libraries and use cases. The team works on a balance of new features / distributed libraries, test infra improvements, debugging, and longer-term architectural improvements to Ray. A snapshot of projects you can work on: - Optimizing performance of large-scale workloads on Ray - Stability and stress testing infrastructure - Improving fault tolerance (HA) As part of this role, you will: Develop high quality open source software to simplify distributed programming (Ray) Identify, implement, and evaluate architectural improvements to Ray core Improve the testing process for Ray to make re

restmachine learningai
View job →
S
Synthesia
📍 London• Full-time
1mo ago

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. We’re looking for an Engineer to join the ML Platform team at Synthesia. Our team builds and operates the systems that allow researchers and product teams to train, serve, and deploy generative models reliably and efficiently . This includes research infrastructure, production serving systems, internal tooling, and the platform interfaces that connect them. A growing part of our mission is making these systems more automation-friendly and agent-oriented , so that workflows can increasingly be operated through reliable tooling rather than manual effort. We’re looking for a strong generalist with a systems mindset: someone who is comfortable working across infrastructure, backend systems, and tooling, and who has seen ML systems in practice. this is not a pure ML Engineer role. We’re especially interested in people who think deeply about reliability, scalability, performance, and resource efficiency in complex production environments. This is a hands-on IC role with significant ownership. You’ll help shape how our ML platform evolves as we scale the number of models, workloads, tools and teams relying on it. What you’ll do Design and improve the platform systems that support model training, evaluation, and product

pythonkubernetesgit
View job →
M
1mo ago

What we're building Mutiny is the self-improving AI infrastructure for GTM teams to execute faster and close more revenue. Our ambition is to do for revenue velocity what Cursor and Claude Code did for engineering velocity. With Mutiny, everyone in sales and marketing gets a bench of GTM athletes that handle any work across their revenue motion and learn from what's actually moved their deals. In April we re-launched the product as an agent-first platform. Anthropic showcased us as a leader in AI GTM. MRR is growing more than 70% month-over-month, with customers like Uber, Rippling, and Snowflake. We're backed by Sequoia, YC, and Insight, and we're building a generational company. The opportunity Most engineers spend their career making predictable systems faster. You'll spend yours making non-deterministic ones trustworthy. As a senior engineer on our AI product team, you'll architect the Campaign Builder and Agent experiences marketers and sellers open every day to go from idea to personalized assets in minutes. You'll partner directly with product, design, and the founders to define what an agent-first GTM platform should feel like, and your calls on architecture, evals, and guardrails compound across thousands of customer accounts. This role is in person in New York City, five days a week, and we ship weekly. What you'll own The core agent surfaces. Architect and ship the Campaign Builder and Agent experiences end-to-end. Frontend, backend, prompts, evals, the whole stack. Reliability on top of LLMs. Make non-deterministic models feel deterministic at the surface. Build the retries, fallbacks, and orchestration so the customer never sees the failure mode. Evals and guardrails. Define how we measure quality, catch regressions, and keep brand and tone consistent across thousands of customer accounts. Speed and feel. AI products live or die by latency and the loop between intent and output. You'll obsess over both, and use coding agents and agent networks to ship f

typescriptpythonai
View job →

About The Role & Team We’re looking for a Fullstack Engineer to join the Core Analytics team, the group behind our flagship analytics product loved by thousands of product teams worldwide. You’ll work with our Product Manager and Designer to define, build, and ship end-to-end features that elevate the Analytics experience. You’ll own your impact from shaping product direction to delivering performant, reliable, and delightful user experiences. Our engineers are customer-focused and solve complex data problems to deliver a better user experience. We deliver value quickly and iteratively. If you’re passionate about building exceptional data-powered experiences that help businesses understand their customers, we’d love to meet you. The team’s mission is to deliver a fast, intuitive, and reliable Analytics platform that empowers customers to confidently uncover insights about their products while driving infrastructure initiatives that strengthen and scale the platform. As a Staff Engineer, you will: Work with Product Management and Design to generate and turn novel ideas for solving customer problems into engineering solutions Drive business impact through owning the end-to-end delivery of projects Have a strong critical thinking and problem-solving mindset with attention to detail Get involved in performance optimization and scaling efforts Drive business impact through leading the highest-leverage projects Have interest or experience in technical leadership of an engineering team Mentor and contribute to the success of other engineers on the team You'll be a great addition to the team if you: 7+ years of experience building and improving robust, scalable backends and developing interactive, user-facing applications or websites. Understand the whole stack and flow of user-facing web applications. Backend experience with application backends, microservices, and supporting high-throughput ingestion systems Strong critical thinking and problem-solving mindset with at

javascripttypescriptpython
View job →
G
Glide
📍 United States• Full-time
1mo ago

👋 Welcome to Glide! At Glide we’re reimagining the banking experience for the modern world . Our embedded fintech platform empowers legacy financial institutions, like community banks and credit unions, to pioneer novel digital experiences for their customers. You’ll be joining an all-star team with engineering, product, and growth experience from Stripe, Google, and Amazon. We’re looking for a talented Fullstack Software Engineer to help us build our initial product. We’re bringing a new perspective to the decades-old financial world , and we’re hoping you can help us do that! Your Responsibilities Lead and mentor a team of engineers, fostering growth, collaboration, and technical excellence. Partner closely with product and design to define requirements, prioritize work, and translate business needs into scalable technical solutions. Oversee development across frontend and backend, ensuring well-tested, secure, and performant code. Establish and maintain best practices in engineering, including CI/CD, test automation, and code review. Provide long-term technical vision for the evolution of Glide’s platform and infrastructure. Help recruit and build a diverse, world-class engineering team. Need-to-Haves Proven experience leading software engineering teams, including mentoring and performance management. Strong technical foundation in fullstack development (JavaScript/TypeScript, React, Node.js, Next.js). Experience designing and maintaining scalable, secure architectures. Deep knowledge of modern frontend technologies (responsive HTML/CSS, state/data fetching libraries like React Query/TanStack or tRPC). Familiarity with cloud infrastructure and DevOps practices (AWS, Docker, Git, CI/CD). Excellent understanding of software engineering best practices: architecture, testing, and security. Strong communication and collaboration skills, with the ability to partner effectively across functions. Nice-to-Haves 7+ years of professional software engineering experience, wi

javascripttypescriptjava
View job →
🔔

Get new back end td reliability lab manager jobs by email

Daily job updates · Unsubscribe anytime