Datadog's Application Performance Monitoring (APM) provides deep visibility into the health, performance, and lifecycle of modern distributed applications, tracing requests from end-user devices (web and mobile) through to backend services. Our goal is to help customers detect root causes faster, optimize application performance, and improve resource efficiency at scale. As the Engineering Manager for APM Serverless, you will help define and deliver the end-to-end serverless APM experience, from auto-instrumentation through troubleshooting, and ensure that OpenTelemetry and Datadog-native customers alike have a frictionless and performant journey. You will also lead efforts to expand coverage of cloud-managed services across providers, ensuring customers can seamlessly trace and monitor critical services in all major and emerging cloud environments. We’re looking for an experienced engineering leader who thrives at the intersection of infrastructure and developer experience. You should care about well-designed APIs, observability-first thinking, and building systems that empower other developers. This is a high-leverage role that will influence how developers across the industry understand and instrument their serverless workloads. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Lead a polyglot team of 8-9 engineers and partner closely with Product and Engineering teams across Datadog to deliver industry-leading serverless capabilities that power consistent, scalable, and intuitive instrumentation across languages. Drive a domain that is technically rich: Lambda, Azure Functions, GCP, OTel billing, Rust, durable functions, distributed tracing across managed services. Engineers on this team work
Jobiba hiring network
Back End Td Reliability Lab Manager Jobs
1,726 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current back end td reliability lab manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Come join and lead the Server Ingress Security team, where we are rearchitecting MongoDB Server’s ingress networking to make MongoDB clusters even more secure. This new team is building the Atlas Network Protection layer, a set of performant, security-critical services that harden MongoDB's pre-authentication attack surface and provides the ability to respond rapidly to emergent threats. We are looking for a talented Lead Engineer to join the team and be founding members, where you will play a crucial role in our multi-year roadmap. Our team champions a strong culture of inclusivity, diversity, and collaboration, and lives MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. If you want to lead a fast-growing team that applies security and systems engineering fundamentals to protect a popular database at scale, join us! We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Candidate Profile 3+ years of experience managing a team of software engineers, including hiring, performance and growth management, compensation planning, and mentoring You have 8+ years of experience building production-quality systems software with large backend/compiled codebases, ideally in Rust. Bonus points for experience with performance profiling, network protocols, TLS, and connection management You have strong technical judgment that you use to effectively guide engineering decisions in security-sensitive or networking-adjacent domains You put the customer first and don't hesitate to cross team boundaries in search of the right solution Solid experience in designing, writing, testing, maintaining, and operating mission-critical software systems Bonus points Professional or advanced academic expertise in the domains of security or networking You enjoy coaching, career development, and creating growth opportunities to help your
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe
Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Full-Stack Engineer to join our Data Acquisition team to build and optimize the interfaces and tools that power our data infrastructure. Responsibilities: Develop and maintain full-stack applications that support data acquisition, including internal tools and dashboards. Collaborate closely with cross-functional teams, including Data Processing, Architecture, and Scaling, to ensure seamless data ingestion and workflow management. Design and implement APIs to facilitate data interactions between internal services and external data sources. Enhance user experience by developing intuitive web-based interfaces for managing and monitoring data pipelines. Optimize backend services for performance, scalability, and security in a distributed computing environment. Work with legal and compliance teams to ensure our data acquisition processes adhere to privacy regulations and best practices. Deploy and maintain infrastructure using Kubernetes and Infrastructure-as-Code (IaC) methodologies. Analyze system performance, conduct experiments, and improve data workflows to maximize efficiency. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in full-stack development. Proficiency in frontend frameworks (React, Vue, or similar) and backend technologies such as Python, Node.js, or Go. Strong expertise in RESTful APIs, GraphQL, and database design (SQL and NoSQL). Experience building data-intensive applications that handle large-scale datasets. Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker). Prior experience with web crawling and large-scale data processing is a
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Platform Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning/sharding strat
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Role Summary At CVS Health®, you’ll be working with a team of passionate colleagues who care deeply, innovate with purpose, hold themselves accountable and prioritize safety and quality in everything we do. Do you enjoy innovation while having fun doing it? If the answer yes, then this role might be for you! Join us and be part of something bigger, innovative and simplification in healthcare. We are seeking a highly experienced and innovative Principal (Director Level) Software Development Engineer to lead the application architecture, design, development, delivery of next-generation digital applications (including Reporting and financial solutions), and optimization of scalable, secure, and high-performance solutions leveraging AI across all major cloud platforms (AWS, Azure, and GCP). This role requires deep technical expertise, AI-enabled solutions, strategic thinking, scalable digital platforms, and enterprise integrations that power critical healthcare and pharmacy experiences and a passion for driving excellence in software engineering practices. This is a senior technical leadership role for a hands-on engineer who can operate across the full stack—from intuitive front-end applications to resilient backend services—while setting architectural direction, influencing engineering standards, and mentoring teams. The ideal candidate combines deep technical expertise, platform thin
As a Software Solution Architect, NVIS at NVIDIA, you will lead the transformation of AI infrastructure. Our NVIS team focuses on developing the next generation of NVIS Central, an agentic software platform with tools, services, and AI agents that automate, simplify, and speed up the work of our delivery organization. This role offers an outstanding chance to create and build LLM-powered agents that improve execution visibility, cut down manual tasks, and standardize workflows. These efforts allow NVIS to grow quickly and with high quality. Join us to bring up, validate, optimize, and upgrade large-scale AI Factory infrastructure for some of the world’s most advanced accelerated computing environments! What you'll be doing: Compose, build, and productionize agentic AI solutions, tools, and applications for the NVIS delivery organization. Develop LLM-based agents, skills, tool-calling workflows, orchestration logic, backend services, APIs, data pipelines, and automation features as part of NVIS Central. Translate field, delivery, operations, and product needs into clear technical builds, agent workflows, and working software. Develop agents that can reason across project data, knowledge bases, operational systems, logs, reports, and delivery workflows. Build workflows that help NVIS teams identify risks, summarize project status, automate repetitive tasks, improve readiness visibility, and simplify handoffs. Work with timely engineering, retrieval-augmented generation, context management, agent memory, function calling, evaluations, and guardrails to build reliable AI systems. Integrate LLMs and agents with internal systems, project data sources, knowledge repositories, reporting tools, and operational workflows. Collaborate closely with software developers, architects, product managers, DevOps/SRE, and NVIS field teams to
**English version below** Doit être local à Montréal Vous souhaitez travailler dans le domaine de la technologie au sein d'une banque d'investissement? Nous recherchons une personne souhaitant rejoindre une équipe dynamique en tant que Développeur Go pour l’un de nos clients. Ce poste est idéal pour un ingénieur backend qui apprécie de travailler à la fois sur le développement logiciel et les opérations. Vous jouerez un rôle clé dans le développement de services backend évolutifs, la conception et la maintenance d’API, ainsi que le soutien aux initiatives d’infrastructure cloud native. Dans le cadre d’une initiative stratégique visant à renforcer la résilience des plateformes, vous contribuerez à la fiabilité et à l’évolutivité des systèmes backend essentiels tout en favorisant l’automatisation et l’excellence opérationnelle. À propos de mtrois : Depuis 2010, mtrois aide ses clients à résoudre leurs défis commerciaux et technologiques. Nous sommes une société de conseil en technologie et en affaires avec une main-d'œuvre mondiale qui réalise des projets commerciaux et informatiques significatifs dans certaines des plus grandes organisations de services financiers du monde. Services principaux Consulting et Conseil Services gérés Programme de diplômés Alumni Programme Alumni Pro Nous avons une présence mondiale et sommes experts dans la fourniture d'une qualité exceptionnelle à notre base de clients, offrant des services de conseil dans les domaines du risque, de la réglementation et de la conformité ; Produits des fournisseurs ; Support d'application ; Développement d'application ; Cyber et sécurité de l'information ; Science des données et DevOps. Notre programme Expert offre aux professionnels expérimentés l'accès à des rôles de premier plan dans la technologie, la finance, l'aviation et l'assurance. Rejoignez-nous pour travailler sur des projets technologiques révolutionnaires, des plateformes de trading internationales aux applications critiques pour les pr
100% Remote | Senior Frontend Engineer | Fintech SaaS Firm About the Role We’re looking for a Senior Frontend Engineer to build and maintain scalable, high-performance user interfaces for our communication platform. You’ll work closely with backend engineers, designers, and product managers to deliver exceptional user experiences while keeping performance, maintainability, and scalability at the core. What You’ll Do Develop and maintain responsive UIs using React JS, TypeScript, JavaScript, HTML5, and CSS. Collaborate with cross-functional teams to design and deliver high-quality features. Write clean, maintainable, and well-documented code. Optimize performance with caching and other best practices. Review code, mentor peers, and uphold coding standards. Debug and troubleshoot production issues promptly. Stay current with frontend trends and bring innovative ideas to the team. Job qualifications: 3–8 years’ experience in web development with a focus on scalability. Expert in React JS, JavaScript, TypeScript, HTML5, and CSS. Strong grasp of responsive design, performance optimization, and client-side session management. Familiarity with Git, CI/CD, and distributed development. Excellent problem-solving and collaboration skills. Preferred/Bonus Skills Experience with React Native or other mobile development frameworks. Familiarity with state management libraries like Redux or Zustand. Experience with modern build tools such as Webpack or Vite. A strong portfolio or active GitHub profile showcasing previous work. Why Join Eltropy? Join a high-impact team building mission-critical backend systems for financial institutions. Work on modern technology stacks in a fast-growing SaaS company. 100% remote work with a collaborative, engineering-led culture. Opportunity to own and influence core backend architecture. About Eltropy Eltropy is a rocket ship FinTech on a mission to disrupt the way people acc
JD : Full Stack with React ,HTML ,CSS 4+ working experience in ReactJs, Html 5, CSS, Responsive design on UI , and Springboot, REST APIs in backend. Experience with Java web application development, Microservices architecture and distributed cloud systems. Experience with REST APIs using Spring boot, Spring data, Spring cloud config, Spring AMQP connector or similar framework in microservices development. Hands on with SQL commands stored procedures. Hands on experience with CI/CD pipelines and setting up various code quality checks in the pipelines. Skills Required UI React JS HTML 5 ,CSS 3 JEST framework Material UI Responsive Design Backend Java/J2ee, Spring MVC, Spring Boot, Micro services Architecture & design Patterns, Docker containers SQL Basic Queries , Stored Procedures AWS EC2/ECS,Lambdas,S3, RDS, SQS
POSITION SUMMARY Full Stack Developer skilled in PHP, Angular, and MariaDB/MySQL to build scalable, secure, and high‑performance web applications. Key Responsibilities Develop and maintain backend services using modern PHP and MariaDB/MySQL. Build RESTful APIs and integrate them with Angular applications. Develop responsive UI components using Angular, TypeScript, HTML5, and CSS/SCSS. Modernize legacy AngularJS components to Angular. Implement authentication/authorization (OAuth, SAML, JWT, MFA) and follow OWASP security practices. Work with Docker, Git/GitHub, and CI/CD pipelines. Monitor and optimize frontend and backend performance (Sentry, CloudWatch). Participate in Agile/Scrum processes and conduct code reviews. EXPERIENCE REQUIRED 4+ years backend experience with PHP. 2+ years frontend experience with Angular (Angular 2+). Strong database design and SQL optimization experience. Experience with REST APIs, RxJS, Docker, and CI/CD. Familiarity with AWS is a plus. QUALIFICATIONS, SKILLS, & KNOWLEDGE Bachelor’s degree in CS/IT or equivalent experience. Strong knowledge of OOP, design patterns, SOLID principles. Proficiency in PHP, MariaDB/MySQL, Angular, TypeScript, HTML5, CSS3/SCSS. Understanding of web security, accessibility, and performance optimization. Experience with testing tools, API documentation, and Agile development. PROFESSIONAL DEVELOPMENT EXPECTATIONS Ability to embrace Clearwater's CLEAR core values (Commitment to Client Success, Lead with Accountability, Integrity & Collaboration, Excellence in All That We Do, Advance Colleague Success, Respect & Transparency) and culture. The base salary range for this role is $25,000-$35,000. Base salary is part of our total rewards package which also includes the opportunity for merit-based salary increases, eligibility for our 401(k) plan, medical, dental, vision, life and disability insurances and leaves provided in line with your work state. Our robust time-off policy
POSITION SUMMARY The Full Stack Developer is skilled in PHP, Angular, and MariaDB/MySQL to build scalable, secure, and high-performance web applications. Key Responsibilities · Develop and maintain backend services using modern PHP and MariaDB/MySQL. · Build RESTful APIs and integrate them with Angular applications. · Develop responsive UI components using Angular, TypeScript, HTML5, and CSS/SCSS. · Modernize legacy AngularJS components to Angular. · Implement authentication/authorization (OAuth, SAML, JWT, MFA) and follow OWASP security practices. · Work with Docker, Git/GitHub, and CI/CD pipelines. · Monitor and optimize frontend and backend performance (Sentry, CloudWatch). · Participate in Agile/Scrum processes and conduct code reviews. SPECIFIC JOB RESPONSIBILITIES 1. Frontend & UI Orchestration · Cross-Framework Mastery: Develop and optimize complex UIs in Angular (v16+) or React , ensuring seamless state management and responsive design. · T3 Stack Architecture: Lead the development of next-gen features using Next.js , Tailwind CSS , and tRPC for end-to-end type safety. · Modern Tooling: Utilize Vite, Turborepo, and modern CSS patterns to maintain a high-velocity developer experience. 2. Full-Stack & Backend Engineering · Type-Safe APIs: Build and consume GraphQL and REST APIs. Leverage tRPC to eliminate the need for manual API documentation between the frontend and backend. · Polyglot Data Management: * Architect relational schemas and optimize complex queries in PostgreSQL . o Manage unstructured security telemetry and real-time logs in MongoDB . · ORM & Migrations: Use Prisma or Drizzle to manage database interactions with 100% type coverage. 3. AI Integration & Innovation · LLM Implementation: Integrate AI m
POSITION SUMMARY We are looking for a Full Stack Developer skilled in PHP, Angular, and MariaDB/MySQL to build scalable, secure, and high-performance web applications. Key Responsibilities · Develop and maintain backend services using modern PHP and MariaDB/MySQL. · Build RESTful APIs and integrate them with Angular applications. · Develop responsive UI components using Angular, TypeScript, HTML5, and CSS/SCSS. · Modernize legacy AngularJS components to Angular. · Implement authentication/authorization (OAuth, SAML, JWT, MFA) and follow OWASP security practices. · Work with Docker, Git/GitHub, and CI/CD pipelines. · Monitor and optimize frontend and backend performance (Sentry, CloudWatch). · Participate in Agile/Scrum processes and conduct code reviews. SPECIFIC JOB RESPONSIBILITIES 1. Frontend & UI Orchestration · Cross-Framework Mastery: Develop and optimize complex UIs in Angular (v16+) or React , ensuring seamless state management and responsive design. · T3 Stack Architecture: Lead the development of next-gen features using Next.js , Tailwind CSS , and tRPC for end-to-end type safety. · Modern Tooling: Utilize Vite, Turborepo, and modern CSS patterns to maintain a high-velocity developer experience. 2. Full-Stack & Backend Engineering · Type-Safe APIs: Build and consume GraphQL and REST APIs. Leverage tRPC to eliminate the need for manual API documentation between the frontend and backend. · Polyglot Data Management: * Architect relational schemas and optimize complex queries in PostgreSQL . o Manage unstructured security telemetry and real-time logs in MongoDB . · ORM & Migrations: Use Prisma or Drizzle to manage database interactions with 100% type coverage. 3. AI Integration & Innovation · LLM Implementation: Integr
Get new back end td reliability lab manager jobs by email
Daily job updates · Unsubscribe anytime