Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr. Staff Network Engineer to join our team. This is a fully remote capacity in the Netherlands role, reporting to the Manager, Network Engineering in the Cloud Ops - Network Engineering department. This position involves a mix of infrastructure, project management, and network engineering responsibilities. You will join our global team as an integral member, taking ownership of the deployment, monitoring, and ongoing operation of our worldwide production infrastructure across diverse data center locations. We are seeking a high-trust collaborator with a background in large-scale enterprise or telecom networks who approaches challenges with a growth mindset and a commitment to achieving engineering excellence. This role requires active participation in our technical on-call rotations, which include support during weekends and holidays to ensure continuous service reliability. What you’ll d
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr Production Engineer to join our team. This role is available as a hybrid opportunity 3 days a week in San Jose, CA or Bellevue, WA reporting to Production Engineering in the Cloud Infrastructure & Operations department. Join Zscaler to be a force multiplier for the reliability of a global platform processing 200+ billion transactions daily across tens of millions of enterprise users. In this role, you will provide the technical vision and hands-on execution to drive an "automation-first" culture across the company. By maturing our observability and architectural standards, you will directly reduce our Mean Time to Mitigate (MTTM) and shape the scalability of our globally distributed, multi-cloud infrastructure. What you’ll do (Role Expectations) Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments Drive an "automation-first" c
Overview We are seeking a hands-on Director of AI Software Engineering to lead and scale AI engineering efforts supporting multiple business units across Governance, Risk, and Compliance (GRC). This role sits at the intersection of product delivery, platform evolution, and applied AI—driving real-world impact across core workflows. This is not a pure management role. We are looking for a builder who leads from the front, someone who has recently written production code, shipped systems end-to-end, and can operate comfortably in ambiguity while aligning teams and stakeholders. What You’ll Do Lead AI Engineering Across GRC Own delivery of AI-powered capabilities embedded directly into business unit workflows (e.g., risk analysis, compliance automation, reporting, due diligence) Partner with product, data, and platform teams to translate business problems into scalable AI systems Stay Hands-On Contribute to architecture, code reviews, and critical path implementation Prototype and validate new approaches (LLMs, agents, retrieval systems, classification pipelines, etc.) Set engineering standards for performance, reliability, and cost efficiency Build and Scale Teams Lead and mentor a high-performing team of AI/ML and software engineers Drive hiring, coaching, and career development Establish a culture of ownership, speed, and technical excellence Drive Execution Deliver production-grade systems—not experiments Balance speed with rigor (security, privacy, compliance) Operate across multiple concurrent initiatives with clear prioritization Communicate and Influence Act as a bridge between engineering and business stakeholders Clearly articulate trade-offs, risks, and outcomes to senior leadership Align cross-functional teams around shared goals and timelines What We’re Looking For Proven Builder 10+ years in software engineering, with recent hands-on coding experience Demonstrated track record of shipping production systems at scale Experience with modern
About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last-mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We're looking for an experienced technical operator to lead live testing, deployment, and operational validation of cutting-edge autonomous technologies. This role sits at the intersection of engineering and operations, helping ensure new capabilities are safely deployed, thoroughly evaluated, and translated into actionable engineering feedback. You’re excited about this opportunity because you will… Lead and oversee live testing across transport, deployment, and safety validation. Partner closely with hardware, software, and business operations teams to validate new product capabilities while providing guidance and mentorship to junior team members. Conduct and document complex tests for autonomous technologies, evaluating robot behavior, identifying issues, and validating new features and requirements. Provide actionable technical feedback to engineering teams based on test outcomes. Exercise technical judgment during live testing by evaluating robot behavior, assessing operational risk, distinguishing expected behavior from product defects, and determining when engineering escalation or additional validation is required. Develop and implement testing processes, protocols, and checklists that improve the safety, efficiency, and reliability of new products. Utilize internal tools to analyze logs, investigate issues, document findings, and track issues through resolution. Mentor junior team members in structured debugging and documentation practices. Conduct detailed analyses and generate comprehensive reports that identify trends, summarize findings, and provide strategic recommendations to engineering and operations partners. We’re excited about you because… 2+ years of exper
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. This is a contract position through our staffing partner Magnit. 108,000.00 - 135,000.00 - 162,000.00 CAD Annual This role is not eligible for the Okta-sponsored benefits listed below. Magnit will provide any locally required benefits. Okta seeks a skilled Senior Recruiter to drive full-lifecycle recruitment and build strategic talent pipelines for our Engineering organization across North America. As part of our AMER Tech Recruiting team, you will be a trusted talent advisor responsible for sourcing, engaging, and delivering top-tier engineering talent while maintaining an "always recruiting" mindset in a fast-paced, high-growth environment. What You'll Be Doing Own full-lifecycle recruitment for Engineering and technical roles (Software Engineering, Site Reliability, Security, TPM, Product) across US & Canada, managing a flexible req load that scales with business priorities. Partner strategically with hiring managers and leadership to understand talent needs, define role scope, advise on talent gap mitigation, and challenge assumptions to ensure hiring decisions strengthen long-term organizational capability. Build and execute talent strategies that balance external hiring with internal mobility, creating sustainable pipelines that reflect commitment to diversity, inclusion, and high-performing engineering culture. Drive metrics-informed recruiting decisions by developing KPIs, analyzing recruiting data, and using insights to optimi
Coordination Systems provides foundational distributed systems building blocks for internal Datadog platforms. Our services cover sharding, consensus, resource protection, configuration distribution, and much more. We are looking for a manager to lead the Coordination Systems - Storage team. This team provides essential configuration storage and distribution systems that are depended upon by almost every service and pod at Datadog. We power critical runtime configuration (e.g. feature flags), complex control planes (e.g. dynamic sharding configuration), and much more. Storage is one of four subteams within Coordination Systems. If successful, the candidate will have opportunities to lead other growing and impactful areas such as Resource Protection. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here/max 6 bullets) Lead a core team of 5 engineers (distributed, with majority in NYC) Lead ceremonies, prioritize and delegate project Stay hands-on with the code, e.g. isolated features, small remediations, investigation follow ups Stay actively involved in operations, incidents, root cause analysis, etc. Constantly promote a culture of operational excellence, organizing gamedays, conducting operational reviews, staying proactive with reliability Who You Are: (Describe role qualifications here/max 6 bullets) Strong distributed systems skills, able to understand and account for a variety of failure modes, well-versed in end-to-end o11y, validation testing, simulation setup, etc. Worked on platform teams before, providing critical infrastructure to internal stakeholders Experienced in handling significant incidents, both as a responder and follow-up ow
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity Are you ready to step into a pivotal leadership role where your engineering depth directly shapes the future of our core platform? As our new Engineering Manager, you will lead a talented, distributed team across US and EU time zones, acting as the critical manager bridging regional collaboration. Our Cloud Foundation team is the backbone of the New Relic platform. In this role, you won't just manage tasks; you will mentor and empower engineers, transitioning our operational framework from a reactive state to a culture of proactive ownership and engineering excellence. You will oversee critical global initiatives, including major regional expansions into FedRAMP High / IL4, India, and Australia. If you thrive on solving complex multi-cloud challenges at an exabyte scale while helping engineers grow in their careers, this is your opportunity to make a lasting impact. What you'll do Empower & Mentor: Lead and nurture a high-performing engineering team across the US and EU, facilitating career development, performance growth, and a collaborative team culture. Drive Strategic Ownership: Champion a shift from reactive delivery to proactive technical ownership, establishing best practices for platform reliability and cross-regional alignment. Lead Regional Expansions: Architect and execute key global infrastructure expansions across complex environments (including FedRAMP High / IL4, India, and Australia). Architect for Extreme Scale: Guide decisions around micr
Who We Are At Justworks, you’ll enjoy a welcoming and casual environment, great benefits, wellness program offerings, company retreats, and the ability to interact with and learn from leaders in the startup community. We work hard and care about our most prized asset - our people. We’re helping businesses get off the ground by enabling them to focus on running their business. We solve HR issues. We’re data-driven and never stop iterating. If you’d like to work in a supportive, entrepreneurial environment, are interested in building something meaningful and having fun while doing it, we’d love to hear from you. We're united by shared goals and shared motivations at Justworks. These are best summed up in our company values, which are reflected in our product and in our team. Our Values If this sounds like you, you’ll fit right in. Department Technology, Data & AI About The Team The Data & Analytics team is Justworks' data and AI center of excellence — a connected set of functions spanning business intelligence, data science, analytics engineering, platform data/ML engineering and AI builders. We've grown significantly over the last several years and continue to build momentum across the business, leveling up maturity in how we deliver insights, govern data, and enable the company to make better decisions. You'll find a highly collaborative & diverse group motivated by craft, shared learning, and a genuine commitment to growing our impact for customers while evolving how we work together for our people. Come build with us! The Data Platform Engineering team sits at the core of this foundation, owning the pipelines, platform infrastructure, and reliability that every other DNA function depends on. Who You Are You're a hands-on engineering leader with deep experience building and operating modern data platforms. You've led teams through scale — evolving architecture, improving reliability, and making hard tradeoffs between speed, cost, and quality. You're as
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Staff, Business Intelligence Engineer, GTM Data Science & Analytics. About the job This position is needed to advance the quality, reliability, and strategic value of Twilio’s Go-To-Market (GTM) data by designing and maintaining robust business data models to enable trusted sales analytics, reporting, data science, & AI capabilities. The Staff, Business Intelligence Engineer will play a mission-critical role in our evolution from traditional BI to AI, ensuring data consistency and empowering data-driven decision-making. Responsibilities In this role, you’ll perform: Data Analysis and Interpretation: Utilize advanced analytical techniques to analyze large datasets, extract meaningful insights, and provide actionable recommendations to enhance revenue operations efficiency Data Monitoring: Develop and maintain automated data reconciliation and quality checks to proactively identify and resolve discrepanci
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Staff, Business Intelligence Engineer, GTM Data Science & Analytics. About the job This position is needed to advance the quality, reliability, and strategic value of Twilio’s Go-To-Market (GTM) data by designing and maintaining robust business data models to enable trusted sales analytics, reporting, data science, & AI capabilities. The Staff, Business Intelligence Engineer will play a mission-critical role in our evolution from traditional BI to AI, ensuring data consistency and empowering data-driven decision-making. Responsibilities In this role, you’ll perform: Data Analysis and Interpretation: Utilize advanced analytical techniques to analyze large datasets, extract meaningful insights, and provide actionable recommendations to enhance revenue operations efficiency Data Monitoring: Develop and maintain automated data reconciliation and quality checks to proactively identify and resolve discrepanci
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Staff, Analytics Engineer, GTM Data Science & Analytics. About the job This position is needed to advance the quality, reliability, and strategic value of Twilio’s Go-To-Market (GTM) data by designing and maintaining a robust business data layer to enable trusted sales analytics, reporting, data science, & AI capabilities. The Staff, Analytics Engineer will play a mission-critical role in driving cross-functional alignment within GTM Data Science & Analytics and Sales Operations on key Sales & Finance business metrics, ensuring data consistency and empowering data-driven decision-making. Responsibilities In this role, you’ll: Design and implement a formal business data layer in dbt, ensuring that data from multiple sources is centralized, reconciled, and serves as a single source of truth for GTM metrics. Collaborate with stakeholders within the GTM Operations team and Finance to define, document, and mai
You will lead a small, hands-on engineering team building the secure, scalable Core Analytics Data Access Platform that accelerates Datadog’s Applied AI and analytics capabilities. The team owns the Data Access Platform — a unified interface that lets AI and analytics teams discover and self-serve production-ready datasets while abstracting underlying systems and embedding required legal and compliance guardrails. In this role you’ll own technical direction, contribute to design and code, and partner closely with Applied AI, Product Analytics, and internal platform teams to provide reliable datasets and APIs for model training and analysis. This role balances day-to-day engineering leadership with long-term platform planning. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead a Hands-On Engineering Team: Manage, mentor, and grow a small team of 2–4 data engineers (mix of senior and junior) across Paris and NYC, fostering technical excellence and career development. Own Technical Direction and Delivery: Define architecture, engineering priorities, and the team roadmap for the Data Access Platform, driving implementation of scalable, secure data pipelines and platform services. Contribute to Design and Code: Spend substantial time coding, reviewing, and shipping critical platform components to ensure performance, reliability, and operational excellence. Partner with Internal Stakeholders: Work closely with Applied AI, Internal Product Analytics, product managers, and platform teams to define data contracts, APIs, SLAs, observability, and curated analytical datasets. Ensure Data Security, Governance, and Reliability: Implement access controls, lineage, monitoring, and compliance guardrails to support safe model training and repeatable analytics workflows.
Manager I, Engineering - Change Experience Platform The Change Experience Platform team builds the internal experiences and platform capabilities that help Datadogs understand, author, route, and safely manage infrastructure changes. The team owns internal UI and CLI frameworks, change-management user experiences, notification and subscription platforms, and infrastructure governance signals used across Datadog’s engineering organization. Its work sits at the intersection of developer experience, infrastructure operations, product design, and change safety. We’re looking for a hands-on technical leader to manage and grow a team of engineers working on the systems that shape how Datadog engineers interact with infrastructure change. You will partner closely with infrastructure, developer experience, platform engineering, and product teams to build reusable interfaces, workflows, and safety mechanisms that make complex change processes easier to understand and safer to execute. This is a high-impact role for someone who enjoys combining product thinking with strong engineering judgment. You will help the team balance framework ownership, platform reliability, internal customer needs, and long-term technical direction across a portfolio that includes UI systems, CLI authoring and publishing, change-management workflows, notification routing, subscriptions, and infrastructure cordon management. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Lead and grow a small team of engineers responsible for internal platforms and product experiences used across Datadog engineering. Help define what “good” looks like for internal developer-facing platforms, including usability, reliability, documentation, adoption, and supportability. Se
Applied AI is where Datadog's ambitious AI bets get built and shipped ( Bits Chat , updog ). We sit at the intersection of research and product: turning promising capabilities from Datadog AI Research lab and the research community into production systems that reach real customers. The team builds the foundations for agentic systems capable of operating at scale in complex production environments. Current bets span agents that run autonomously at scale, context and memory layers that make those agents more intelligent over time, and tools that help customers build and validate AI-native services in production. The mandate is to move fast from idea to customer impact, and when a product finds its footing, to set it up for growth. As an Engineering Manager I in Applied AI, you will lead a team of engineers and applied scientists working on one of these challenges. You will define technical direction, run short feedback loops, make deliberate decisions about what to pursue or stop, and work closely with product managers, research teams, and cross-functional partners to ship AI capabilities that matter. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Lead and develop a team of engineers and applied scientists focused on building the foundations for agents operating at scale Work closely with product managers, research teams, and cross-functional partners to shape the team's bets from initial framing through to broader adoption, with a clear definition of success criteria at each stage Own end-to-end delivery of high-quality AI systems, from early research exploration to production-grade reliability, with high standards for operational excellence, system reliability, and technical quality Navigate the unique challenges of shipping AI-powered products: balancing quali
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime