We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Data, Identity, & API Platform group at New Relic builds the foundation for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the proprietary telemetry database technology at the core of our business. We own our software from top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally be on-call to respond to system failures; so we prioritize work that keeps the lights on and the pager quiet, in addition to the work that powers all of our new products and streams of data. If the idea of working on systems that process millions of messages per second and handle petabytes of data excites you, then you may be an excellent fit! What you'll do Own the New Relic query language and gateway stack Proactively participate in cross-functional committees to move the query language and gateway forward, ranging from collaborations with AI, Visualizations, and Data Processi
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity As a Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with a proven track record of building and scaling resilient infrastructure. What you'll do Platform Orchestration: Work on a large-scale K8s infrastructure platform, ensuring high availability and performance. Automation: Drive the evolution of internal tooling to streamline platform delivery. This role requires Experience: Solid hands-on background in DevOps, Site Reliability, or Infrastructure Engineering. Kubernetes Mastery: Deep internal knowledge of K8s primitives (Deployments, StatefulSets, Services) and hands-on experience writing custom Kubernetes Operators. Golang Proficiency: Proficiency in Go, specifically for infrastructure automation and systems programming. Operations-Heavy Mindset: A proven track record of managing production environments and handling high-severity incidents. Cloud Infrastructure: Hands-on experience with cloud-native scaling tools (e.g., Karpenter, Cluster API) and Day 1/Day 2 operations of K8s clusters. Tooling: Familiarity with Helm and GitOps workflows (e.g., ArgoCD or Flux). Please note that visa sponsorship is not available for this position. Fostering a diverse, welcoming and inclusive environment is important to us. We work hard to make everyone feel comfortable bringing their best, most authentic selves to work every day. We cele
#TeamNextdoor Nextdoor (NYSE: NXDR) is the essential neighborhood network. Neighbors, public agencies, and businesses use Nextdoor to connect around local information that matters in more than 350,000 neighborhoods across 11 countries. Nextdoor builds innovative technology to foster local community, share important news, and create neighborhood connections at scale. Download the app and join the neighborhood at nextdoor.com . Meet Your Future Neighbors As a Software Engineer at Nextdoor, you’ll join a focused team of developers, product managers, and designers who are passionate about using technology to cultivate a kinder world where everyone has a neighbor they can rely on. We are a small, high performing team of engineers that wear multiple hats and prioritize impact for our customers. We care about moving fast and delivering impact, without compromising on quality and reliability as guided by our Engineering Principles . You will have the opportunity to learn from your co-workers and teach them. As a team, we will make each other better and build great software. At Nextdoor, we operate in an AI-first environment and expect every team member to actively use AI tools as part of their workflow. We aren't looking for prompt engineers; we’re looking for people who use tools like Claude, Gemini, ChatGPT, and Glean to challenge their own thinking and take full ownership of AI-assisted outputs. We also offer a warm and inclusive work environment that embraces a hybrid employment model, blending an in office presence and work from home experience for our valued employees. The hiring team will go over these expectations with you if you are being considered for a role near one of our offices in San Francisco, Los Angeles, Chicago, Dallas, New York, and London. The Impact You’ll Make We believe in empowering our teams to own all aspects of bringing Nextdoor to life. As such, you’ll get the opportunity to make key contributions across our engineering stack - thi
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Devices Platform team's mandate is to lay the foundation of Nuro's onboard software for our sensor and compute platform, including device drivers, inter-device protocols and pipelines, and device runtime APIs. Sensors and compute hardware are the eyes, ears, and brains of our self-driving robots. We are creating the hardware-agnostic platform to be used by the perception and autonomy SW stack, and to realize the full potential of our sensor and compute HW in reliability, quality, and performance. The projects we work on are high impact and high visibility within Nuro. This team is also responsible for working with internal stakeholders and external suppliers to define, evaluate, integrate the next generation HW platform for Nuro's products and to build the necessary tooling to assist continuous testing and validation. About the Work Design and develop sensor and compute systems for robotics Architect and/or deploy Nuro
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Telemetry Data Platform group at New Relic builds the foundation for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the proprietary telemetry database technology at the core of our business. We own our software from top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally be on-call to respond to system failures; so we prioritize work that keeps the lights on and the pager quiet, in addition to the work that powers all of our new products and streams of data. If the idea of working on systems that process millions of messages per second and handle exabytes of data excites you, then you may be an excellent fit! What you'll do Develop new features with a focus on optimizing performance and efficiency Collaborate with the team to implement scalable solutions and enhance application performance Identifying and acting on opportunities to improve the reliability of our services This role requires 2+ years of professional experience in distributed SaaS software development. Proficiency in Java programming, expertise with algorithms and data structures, and building high-throughput software following best-practices. Deeper understanding of distributed systems and their core challenges. Experience using the command line to manage, investigate, and fix things when they’re broken. Expe
#Team Nextdoor Nextdoor (NYSE: NXDR) is the essential neighborhood network. Neighbors, public agencies, and businesses use Nextdoor to connect around local information that matters in more than 350,000 neighborhoods across 11 countries. Nextdoor builds innovative technology to foster local community, share important news, and create neighborhood connections at scale. Download the app and join the neighborhood at nextdoor.com . Meet Your Future Neighbors As a Senior or Staff Software Engineer at Nextdoor, you’ll join a focused team of developers, product managers, and designers who are passionate about using technology to cultivate a kinder world where everyone has a neighbor they can rely on. We are a small, high performing team of engineers that wear multiple hats and prioritize impact for our customers. We care about moving fast and delivering impact, without compromising on quality and reliability as guided by our Engineering Principles . You will have the opportunity to learn from your co-workers and teach them. As a team, we will make each other better and build great software. At Nextdoor, we operate in an AI-first environment and expect every team member to actively use AI tools as part of their workflow. We aren't looking for prompt engineers; we’re looking for people who use tools like Claude, Gemini, ChatGPT, and Glean to challenge their own thinking and take full ownership of AI-assisted outputs. We also offer a warm and inclusive work environment that embraces a hybrid employment model, blending an in office presence and work from home experience for our valued employees. The hiring team will go over these expectations with you if you are being considered for a role near one of our offices in San Francisco, Los Angeles, Chicago, Dallas, New York, and London. The Impact You’ll Make We believe in empowering our teams to own all aspects of bringing Nextdoor to life. As such, you’ll get the opportunity to make key contributions across our enginee
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea
The Team + The Role The Core team builds features and platform capabilities that power critical areas of the Pendo product experience. The team supports high-stakes customer use cases, shapes how systems evolve over time, and creates a more consistent experience across the platform. This work sits in complex product areas where delivery speed, system reliability, and long-term maintainability all matter. As a Sr. Software Engineer, you will own complex problem spaces and drive them forward independently. You will apply strong engineering judgment to ambiguous challenges, partner closely with Product and Design, and stay accountable to outcomes, not just output. You will break large efforts into small, continuously shippable increments, bring others along with you, and leave the codebase better than you found it. This role is based in our Raleigh office. What this looks like day-to-day Feature ownership: Own complex, ambiguous features end-to-end by structuring work into small, independently shippable increments. Drive delivery with clarity, strong judgment, and accountability for customer and system impact. Production engineering: Write production-ready code and define testing approaches based on risk and system impact. Use safe rollout patterns to reduce risk and enable incremental delivery. Incident response: Respond quickly when production issues arise in your area. Resolve issues and drive follow-up improvements that prevent recurrence. AI-enabled development: Use AI tools as a core part of your daily workflow for code generation, debugging, test writing, and task decomposition. Help raise the bar for how the team uses AI in development workflows. Technical collaboration: Deliver actionable code reviews, accurate documentation, and constructive contributions to technical discussions and planning sessions. Help teammates make better decisions through clear, direct input. Continuous improvement: Identify recurring friction points in code, tests, tooling, or proces
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The App Core team is looking for a curious, growth-minded Software Engineer II to work on the core infrastructure of Smartsheet, focusing on the reliability, stability, and scale of our foundational systems. This role is a perfect fit for a developer who has the software engineering basics down and wants to fast-track their skills at the intersection of generalist backend engineering and Infrastructure as Code (IaC). This team practices mob programming for the majority of the workday. We operate through collective code ownership, meaning you'll be collaborating in real-time with teammates most of the time—not working solo on isolated tasks. We focus on building high-quality software while adhering to rigorous operational best practices. If you are passionate about continuous learning and ready to dive into the core foundation of our platform, we want to hear from you. You Will: Contribute to Core Infrastructure: Help build and maintain the core infrastructure that serves as the backbone for Smartsheet. Contribute to a robust environment that ensures the foundational reliability, stability, and performance expected by all of our users. Collaborate via Mob Programming: Work closely with the team every day in a re
About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Engineer on the Detection and Response (D&R) team at Roblox, you’ll protect our user community alongside the underlying platform infrastructure. You’ll design high-fidelity detections, engineer security data platforms, and respond alongside the team during incidents. This is a hybrid in-office role in San Mateo. You Will: Deliver robust D&R capabilities: Engineer high-fidelity detections end-to-end. Lead partners through threat modeling and logging, to deploying actionable alerts, while keeping false positives low. Build security data pipelines: Develop security data pipelines and actively contribute to internal software and data platforms, collaborating across engineering teams. Ensure service reliability: Participate in an on-call rotation to keep detection and response services healthy. Embody security culture: Serve as a trusted security partner across Roblox, helping protect our community and enterprise while fostering a culture grounded in trust, ownership, and shared responsibility. You Have: 3+ years of experience in Security Data Engineering: You have built services that are efficient, reliable, and scalable using programming languages like Golang or Py
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Engineering Manager, Communications, you'll lead the team responsible for in-game text chat — one of the most-used surfaces on Roblox and a cornerstone of how our community connects. Every day this system carries billions of messages across hundreds of millions of users, in real time, across 2D and 3D spaces, on every device we support. You'll own the roadmap and the engineering org behind it: building rich, immersive, and engaging communication experiences that let people and creators express themselves safely and seamlessly — better than in real life. This is a role for a leader who thinks like a product builder as much as an engineer. You'll balance the demands of massive scale and rock-solid reliability with a relentless focus on the user experience, shipping features that make conversation on Roblox feel effortless, expressive, and safe. You'll grow and develop a team of engineers, set technical direction, and partner deeply across product, design, trust & safety, and infrastructure to define what communication on Roblox becomes next. You Will Lead, grow, and develop a team of software engineers building the in-game text chat platform, setting a high bar for engineering
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Sharing team, you will lead large, multi-team initiatives with long-term technical vision and group-level impact. You will be expected to define and drive the platform-wide strategy powering how millions of users capture, share, and discover content on Roblox. In this role, you will set engineering standards, mentor senior engineers, and serve as a key architect of our long-term technical direction. You Will Drive Strategy & Execution: Own the outcome of complex, business-critical programs spanning several teams, often lasting years. Innovate at Scale: Develop and drive a multi-year technical vision for content creation and sharing, anticipating scale, technology, and business evolution. Elevate Reliability: Lead high-severity incident response across groups; drive durable systemic solutions that improve reliability and velocity. Architect Foundations: Regularly improve shared infrastructure and foundational systems, introducing frameworks that uplift development speed across the organization. Align Teams: Aligns multiple teams on shared technical direction, producing detailed design docs, phased roadmaps, and planning models that balance short an
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The platforms in our Engineering Acceleration org dictate how thousands of Roblox engineers ship, and how every backend service gets built, deployed, and kept reliable at scale. They sit on the critical path for production reliability, developer velocity, and security posture. Reporting to Andrew Swerdlow, you will rethink the entire engineering toolchain from the ground up to be agentic-native, creating a world where AI agents are first-class participants in software development and humans set direction and supervise. This is a unique opportunity to define what modern engineering infrastructure looks like at one of the largest platforms in the world. You will: Reimagine the engineering toolchain as agentic-native, designing the platforms, guardrails, and feedback loops that let AI agents safely drive migrations, validation, and routine operational work. Set a bold technical direction for AI-driven quality, including agent-generated tests, automated coverage of untested paths, intelligent verification, and continuous-deployment workflows. Make software quality and SEV prevention a measurable property of the platform by investing in safe-change mechanisms, automated verification, progressive
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver effective performance ads to our users, and more business values to our advertisers. We’re looking for an EM to lead a team of exceptional ML infrastructure engineers, build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You Will: Lead strategic planning and roadmap execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference. Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature). Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment. Recruit, mentor, and grow a high-performing team of ML infrastructure engineers. You Have: 5+ years of experienc
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime