The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde
We are seeking a Staff engineer to design, build, and operate the internal and external Observability stack for the MongoDB platform. Tens of thousands of customers depend on our Observability stack to monitor their database clusters and to generate actionable alerts to safeguard critical workloads. The Collections team is a newly formed team within MongoDB's Observability & Adoption Org focused on making telemetry onboarding and collection significantly easier across MongoDB. We own key parts of the observability collection stack, including onboarding experience, telemetry collection agents across the data and control planes, and ingestion services for metrics, logs, and traces that support both internal and customer observability in MongoDB, driving insights, recommendations, and alerting. Our mission is to reduce friction for teams implementing and iterating on Observability while partnering closely with development teams to instrument their services using shared best practices, helping define the conventions our telemetry should follow, and building collection and ingestion systems that are stable, performant, secure, well-documented, and self-service. We also work closely with the Data Pipeline and Storage & Query teams to help ensure MongoDB has a stable and performant observability stack end to end. This is an opportunity to join a team shaping how observability works across MongoDB and to have outsized impact on both the developer experience and the reliability of the platform underneath it. As MongoDB Atlas and its supporting infrastructure continue to experience rapid growth, the demand for high-cardinality observability data for internal and external use cases means we need to continually innovate and scale our systems to the next level. For example, MongoDB Observability systems need to handle 10’s of billions of metrics time series, all whilst processing petabytes of logs, traces, and events. Our stack includes VictoriaMetrics, Grafana, Sp
We are seeking a Staff engineer to design, build, and operate the internal and external Observability stack for the MongoDB platform. Tens of thousands of customers depend on our Observability stack to monitor their database clusters and to generate actionable alerts to safeguard critical workloads. The Collections team is a newly formed team within MongoDB's Observability & Adoption Org focused on making telemetry onboarding and collection significantly easier across MongoDB. We own key parts of the observability collection stack, including onboarding experience, telemetry collection agents across the data and control planes, and ingestion services for metrics, logs, and traces that support both internal and customer observability in MongoDB, driving insights, recommendations, and alerting. Our mission is to reduce friction for teams implementing and iterating on Observability while partnering closely with development teams to instrument their services using shared best practices, helping define the conventions our telemetry should follow, and building collection and ingestion systems that are stable, performant, secure, well-documented, and self-service. We also work closely with the Data Pipeline and Storage & Query teams to help ensure MongoDB has a stable and performant observability stack end to end. This is an opportunity to join a team shaping how observability works across MongoDB and to have outsized impact on both the developer experience and the reliability of the platform underneath it. As MongoDB Atlas and its supporting infrastructure continue to experience rapid growth, the demand for high-cardinality observability data for internal and external use cases means we need to continually innovate and scale our systems to the next level. For example, MongoDB Observability systems need to handle 10’s of billions of metrics time series, all whilst processing petabytes of logs, traces, and events. Our stack includes VictoriaMetrics, Grafana, Sp
MongoDB is building a world-class team in North America to create tooling that helps customers modernize their applications and migrate their data from legacy relational databases to MongoDB in real-time. As companies modernise legacy workloads and data ecosystems, they are increasingly drawn to the flexibility and scalability of the document model. The tools developed by the Code Generation and Data Migration team are critical in this journey, helping customers with schema modeling, code generation, initial data loads, and continuous data synchronization. We're looking for a Software Engineer with a strong background in computer science fundamentals, systems design, experience in the Java ecosystem, streaming systems, and data-intensive applications to join our engineering team. In this role, you will be instrumental in designing, building, and optimizing the underlying data structures, algorithms, and database interactions that power our generative AI platform, code generation and migration tools. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation and building a sophisticated data migration suite using a modern technology stack, which includes Java, Spring Boot, Kafka, Debezium, and React.You will work on critical components that ensure the scalability, efficiency, and reliability of our services, collaborating closely with AI researchers, product management and other engineers to design and implement cutting-edge products that solve complex customer challenges. This role will be based out of Washington, Oregon, or California. The ideal candidate for this role will have 2+ years of engineering experience in backend systems, distributed systems, or core platform development Experience in one or several of Java, Rust, C/C++, and/or Python, with a strong understanding of systems-level programming, memory ma
We are looking to speak to candidates who are based in Gurugram for our hybrid working model. About the Role We are looking for a Staff Integration Engineer (Workato & API Integration) to join our GTMTech team. This strategic role will lead the design, implementation, and governance of enterprise-grade integrations that power our core business processes across GTM systems, with a primary focus on Workato-based integrations and modern API management patterns. You will own the architecture for critical integration domains such as Quote-to-Cash and other high-impact GTMTech programs, ensuring our integration landscape is scalable, secure, observable, and aligned with best practices for event-driven and API-first designs. You will partner with engineering, architecture, security, and business stakeholders to define standards, mentor other integration engineers, and drive continuous improvement in how we connect systems and data. The GTMTech team is focused on high-impact, large-scale technology programs, such as Quote-to-Cash. Our mission is to enhance company efficiency, boost profitability, and enable data-driven decision-making by streamlining systems, optimizing processes, and minimizing manual, low-value work through automation and AI. Key Responsibilities Lead the end-to-end architecture, design, and implementation of Workato-based integrations and APIs across GTM systems (e.g., Salesforce, NetSuite, HRIS, Google Workspace) with a focus on scalability, reliability, and security Define and evolve integration standards, patterns, and best practices, including canonical integration patterns, error-handling strategies, observability, and operational runbooks Design and review complex, event-driven integration workflows leveraging technologies such as Kafka or equivalent messaging platforms, ensuring robust handling of topics, producers/consumers, durability, and retry mechanisms Drive API-first and MCP-native design for GTM integrations, leveraging RESTful APIs al
The MongoDB Customer Observability Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The team is responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy globally distributed MongoDB clusters in just minutes. We're seeking a Senior Software Engineer to join our team to tackle exciting challenges within the Observability space. You'll contribute to developing tools and platforms that help our users understand the health and performance of their MongoDB deployments. This includes collecting metrics, monitoring slow queries, and offering actionable insights such as index and schema suggestions that improve the speed, efficiency, and overall reliability of their databases. This role provides a unique opportunity to drive engineering excellence across both dimensions of observability, leveraging technologies being developed within the Customer Observability group and contributing directly to the success of our customers and our product teams. If you're passionate about large-scale systems, digging deep into telemetry data, and building tools that make a real impact both internally and externally, we’d love to have you on board! We are looking to speak to candidates who are based in Dublin for our hybrid working model. We're looking for someone who Has at least 5 years of experience as a backend or full stack engineer Enjoys collaboration and being part of a team Is approachable, curious, and intellectually honest Is a backend engineer with a willingness to take on frontend tasks or a full-stack developer with a bias towards backend Has written backend systems in a compiled language (Java, C#, Go, etc.) Has experience with the design and architecture of a modern, scalable web application Enjoys chasing down difficult problems in a distributed environment and on an database diagnostic/operation level Always strives to expand their knowledg
We’re looking for a Software Engineer 3 to help bring Voyage’s embedding models - used for semantic search, retrieval, and AI-native experiences; to the platforms and environments where customers already run their workloads, beyond first-party MongoDB Atlas. You’ll join the broader Search and AI Platform organization and collaborate closely with the engineers building Voyage’s first-party inference. Together, we’re extending that platform across cloud marketplaces, third-party inference providers, and self-managed deployments so customers get the same Voyage models, behaving consistently, wherever they choose to run them. As a Software Engineer 3, you'll focus on building the systems, tooling, and deployment workflows that power third-party model delivery. You'll own key components of how Voyage models are packaged, validated, and deployed, work across teams to ensure tight integration with the core inference platform, and contribute to delivery surfaces designed for reliability, observability, and ease of use. We are looking to speak to candidates who are based in Sydney for our hybrid working model. What you'll do Port and tune the model server that runs Voyage embedding and reranking models: improving inference performance, consistency, and runtime behavior across environments Productionize new Voyage models for delivery beyond first-party Atlas, owning the packaging, configuration, and deployment workflows that get them running on AWS, Azure, GCP and more Design correctness, correlation, and performance validation that proves third-party deployments match first-party behavior Build operability into every surface: structured logging, metrics, diagnostics, and health checks with tools like Prometheus and OpenTelemetry Debug problems that span model servers, containers, deployment configuration, and partner cloud environments Work alongside Voyage's model-serving teams, and partner with GTM, SAs, TSEs, and strategic customers on the hardest external deployments Who
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Product Auth0 is a developer-friendly identity platform that simplifies authentication and authorization for applications. Designed by developers for developers, we make access to applications safe, secure, and seamless for the more than 100 million daily logins around the world. Our modern approach to identity enables this Tier-Ø global service to deliver convenience, privacy, and security so customers can focus on innovation. Know more about our product at https://auth0.com/ . The Role We are seeking a founding Engineering Manager to lead and bootstrap our newest team: Core Frontier . This team sits at the vital intersection of deep product innovation and the actual customer experience. Your mission is to ensure that the sophisticated features developed across the Core Identity organization (such as Native to Web, Cross-App Access and Custom Token Exchange) are translated into a seamless and intuitive journey for our users. You will act as the champion for a complete and coherent product experience, bridging the gap between complex internal logic and the polished final result our customers interact with every day. As the inaugural manager for Core Frontier, you will lead the hiring and onboarding of a high-caliber team in Bengaluru while establishing the operational rhythms that drive success. You will partner with global engineering leaders to maintain our high standards for security and reliability, ensuring that every release meets the rigor
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we’re building the future of secure, enterprise-grade AI Agents . We’re looking for a Principal Engineer to join our global AI Engineering team. In this role, you will be instrumental in designing and building the intelligent, user-facing experiences and the end-to-end AI solutions that power them. This is a senior individual contributor role for a hands-on engineer who can set technical direction, mentor others, and partner closely with our cross-geo counterparts as part of one global team. What you'll do : Drive the architecture and design of AI solutions, leading cross-functional initiatives across Product, Design, and Data Science. Design, build, and refine the core backend services that power our AI solutions, including LLM orchestration, RAG pipelines, and generative AI features. Build high-performance Agentic Experiences (AX) for web and mobile, engineered for streaming responses and low latency. Champion observability and operational excellence to ensure our AI services meet enterprise-grade standards for reliability and performance. Develop robust backend services to power our AI solutions, including LLM orchestration, RAG pipelines, and generative AI features Enable the successful delivery of key AI projects through technical leadership and hands-on execution. Mentor engineers, raising the bar on technical craftsmanship and solution quality across the full stack. Collaborate with cross-geo peers to ensure globally aligned designs, shared
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business-critical online systems at massive scale, driving the reliability, efficiency, and evolution behind every core Pinner and Advertiser experience. You'll lead major efforts like multi-region deployment and Kubernetes migration, set the standard for operational excellence, and define the long-term vision for our online serving infrastructure, supporting machine learning and product innovation across the company. This is an opportunity for high-impact technical leadership, broad visibility, and cross-functional influence at the heart of Pinterest’s platform. What you’ll do: Improve reliability, scalability and infra efficiency for Pinterest’s critical online systems across storage and caching, online service and realtime analytics syste
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is looking for a Principal Engineer to lead the technical strategy and execution for our Indexing & Retrieval Infrastructure across Core, Ads, and Shopping. This role sits at the center of our delivery infrastructure organization and will shape how fresh, relevant, and cost-efficient candidates are generated for the discovery experiences that power Pinterest at scale. What you’ll do: Define and drive the long-term technical vision for indexing and retrieval infrastructure across Core, Ads, and Shopping, aligning architecture investments to measurable improvements in freshness, quality, coverage, reliability, and cost efficiency. Lead cross-org initiatives to modernize Pinterest’s indexing architecture, including advancing unified real-time and incremental retrieval systems that support high-scale, high-quality candidat
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . What you’ll do: Lead a cross-platform team responsible for image rendering and video playback across Pinterest apps and web. Set the team’s technical direction and drive execution on quality, performance, and reliability for media experiences. Partner closely with Media Transcoding, Performance, Data Science, Product and Infrastructure teams to improve how media is delivered and experienced across Pinterest. Drive initiatives to reduce media load times, improve playback smoothness, and ensure high visual quality at scale. Guide the team through platform and architectural decisions, balancing product needs, technical investments, and long-term maintainability. Build and grow a high-performing team through hiring, coaching, and developing engineers. Raise the bar on engineering excellence through strong operational practices, performance mea
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Build Systems team within Figma’s Developer Experience organization owns Figma’s build and CI infrastructure, enabling engineers to ship changes to production quickly and safely. We build and operate core platforms across our polyglot monorepo, including build systems, artifact repositories, merge queues, test frameworks, and CI pipelines. We’re looking for an experienced technical leader to help shape these platforms, uplevel the team, and deliver high-impact platforms that accelerate engineering velocity. The ideal candidate has deep experience with large monolithic codebases, builds durable and scalable systems, and is motivated by solving high-leverage problems that amplify productivity across the engineering organization. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Drive technical roadmap and strategy for the Build Systems team Partner with cross-functional teams and leadership to identify developer pain points and design elegant, scalable solutions Lead complex, multi-quarter initiatives reducing build/test times and improving CI reliability, all while balancing technical excellence with pragmatic delivery Design, build, and maintain modern developer tools including scalable build systems, distributed CI platforms, and test frameworks that serve thousands of engineers Architect and implement large-scale infrastructure on AWS that powers our entire build pipeline to ensure reliability, performance, and cost
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Data Platform team at Figma builds and operates the foundational systems that power analytics, AI/ML, and data-driven decision-making across the company. We serve a diverse set of stakeholders, including AI researchers, machine learning engineers, data scientists, product engineers, and business teams that rely on data for insights and strategy. Our team owns and scales critical platforms such as the Snowflake data warehouse, ML Datalake, orchestration and pipeline infrastructure, and large-scale data ingestion and processing systems, managing all data flowing into and out of these platforms. Despite being a small team, we take on high-scale, high-impact challenges. In the coming years, we're focused on building the data infrastructure layer for Figma's AI-powered products, driving cost and performance optimizations across our data stack, scaling our ingestion and reverse ETL capabilities for new product use cases, and strengthening data quality, reliability, and compliance at every layer. If you're passionate about building scalable, high-performance data platforms that empower teams across Figma, we'd love to hear from you! This is a full-time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Design and build large-scale distributed data systems that power analytics, AI/ML, and business intelligence across Figma. Develop batch and streaming solutions to ensure data is reliable, efficient, and scalable across the company. Manage and evo
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime