Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 Singapore• Full-time• Remote
21 days ago

About the Team OpenAI’s Applied AI Engineering team helps organizations turn frontier AI capabilities into safe, reliable, and high-impact production systems. We work with customer executives, product and engineering teams, security leaders, and transformation teams to identify valuable opportunities, accelerate technical implementation, and scale what works. Enterprise deployments are defined by complexity rather than any one industry: existing architectures, diverse data environments, security and governance requirements, multiple stakeholder groups, and organization-wide change. We turn lessons from these deployments into better products and reusable patterns for customers everywhere. About the Role As an Applied AI Engineer you will partner directly with leading organizations to design, build, and deploy AI systems that deliver measurable business outcomes. You will combine deep technical judgment, hands-on engineering, and customer leadership to take ambitious ideas from use-case selection and architecture through prototyping, evaluation, production launch, and scale. You will write and debug code, build evaluation systems, resolve complex integrations, and guide decisions involving model behavior, reliability, latency, cost, safety, security, governance, and operational readiness. Success is measured by production systems, sustained adoption, and meaningful customer impact—not simply activity or successful demonstrations. This is a rare opportunity to work on consequential real-world deployments at the frontier of AI while directly influencing how OpenAI’s products evolve. This role is based in Singapore. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Partner directly with enterprise customers to identify high-value opportunities and translate them into technical architectures, implementation plans, evaluation strategies, and measurable success criteria. Design, build, and de

REMOTEjavascripttypescriptpython
View job →
C
21 days ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role is for people who love building tools for their coworkers. The Internal Applications team creates tools that help us create better models. In this role you will collaborate with internal stakeholders, which include annotators, ML researchers, product managers and more. Join our team of builders who create tooling that will pave the way for the next generation of large language models! As a Full-Stack Software Engineer on the Internal Applications team, you will: Work with a small talented and enthusiastic team of software engineers Contribute to delightful experiences for our user-facing products, meticulously crafting code for browsers and servers Collaborate and grow with your engineering colleagues of all levels through direct pairing sessions, architectural designs, documentation and talks Identify and remove roadblocks to enable your team to increase its engineering velocity. Build resilient systems that are mission-critical Keep up with the cutting edge and adopt new technologies to improve performance and reliability You may be a good fit if: You have experience shipping products with a large numb

REMOTEtypescriptpythonreact
View job →
O
Okta
📍 Bengaluru• Full-time
21 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the Role Okta is the identity standard. The Okta Identity Cloud is an independent and neutral platform that securely connects the right people to the right technologies at the right time. We help organizations secure and manage their extended enterprise while transforming their customers’ experiences. With thousands of global customers, 7,000+ app integrations, and over 200 million registered users, we are only getting started. As a member of the Developer Productivity Engineering team, you will tackle high-impact challenges across development environments, AI enablement for engineering, scalability, and stability. Grounded in Okta’s core value— Always Secure. Always On. —your work directly powers developer velocity, system reliability, and software quality at scale. You will act as a force multiplier for our engineering teams by identifying workflow bottlenecks, pioneering AI integrations, establishing best practices for code organization, and maintaining performant development environments. What You’ll Do Design & Automation: Build and ship automated solutions that allow developers to deliver features rapidly without compromising quality, stability, or security standards. Environment Performance & Tuning: Analyze local development workflows, build tools to track operational metrics, and profile/tune development environments (including codebase modifications). AI & Tooling Enablement: Leverage AI technologies across the development stack

javaawsazure
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
22 days ago

About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal

REMOTEawsrestai
View job →
M
Mixpanel
📍 New York• Full-time• From $130K/yr
22 days ago

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Marketing team at Mixpanel is responsible for pipeline generation across paid, website, and product-led channels. This role sits within that team as its dedicated engineering owner — accountable for the technical systems that power how people discover, evaluate, and request Mixpanel. You are the only engineer dedicated to this surface, which means you set its technical direction rather than execute against someone else's. You work directly with our marketing teams and partner with Growth Engineering on shared systems. Website design and content are handled by our web design team; your focus is the technical infrastructure, integrations, and systems that sit underneath. About the Role As our Marketing Systems Engineer, you own the technical layer of our marketing engine: the integrations and infrastructure that power our marketing website. The website is our primary lead capture system that turns traffic into pipeline. You own the systems that connect our marketing surface to Salesforce, Customer.io/Hubspot, and our broader GTM stack. You set the technical roadmap for this surface, move fast, measure impact, and treat reliability as a first-class concern rather than a cleanup task.You will also collaborate with Growth Engineering on shared infrastructure including handraiser routing and tracking. Responsibilities Own the technical health of the Mixpanel marketing website: page speed, WCAG compliance, Google Tag Manager, technical SEO, and third-party integrations including Qualified, Optimizely, and TrustArc. Own the integrations between the marketing website and our GTM stack t

javascriptpythonjava
View job →
M
Mixpanel
📍 San Francisco• Full-time• From $130K/yr
22 days ago

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Revenue Marketing team at Mixpanel is responsible for pipeline generation across paid, website, and product-led channels. This role sits within that team as its dedicated engineering owner — accountable for the technical systems that power how people discover, evaluate, and request Mixpanel. You are the only engineer dedicated to this surface, which means you set its technical direction rather than execute against someone else's. You work directly with our marketing teams and partner with Growth Engineering on shared systems. Website design and content are handled by our web design team; your focus is the technical infrastructure, integrations, and systems that sit underneath. About the Role As our Marketing Systems Engineer, you own the technical layer of our marketing engine: the integrations and infrastructure that power our marketing website. The website is our primary lead capture system that turns traffic into pipeline. You own the systems that connect our marketing surface to Salesforce, Customer.io/Hubspot, and our broader GTM stack. You set the technical roadmap for this surface, move fast, measure impact, and treat reliability as a first-class concern rather than a cleanup task.You will also collaborate with Growth Engineering on shared infrastructure including handraiser routing and tracking. Responsibilities Own the technical health of the Mixpanel marketing website: page speed, WCAG compliance, Google Tag Manager, technical SEO, and third-party integrations including Qualified, Optimizely, and TrustArc. Own the integrations between the marketing website and our GTM stack t

javascriptpythonjava
View job →
O
22 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description Okta is the leading independent provider of enterprise identity. The Okta Identity Cloud enables organizations to securely connect the right people to the right technologies at the right time. With over 6,500 pre-built integrations to applications and infrastructure providers, Okta customers can easily and securely use the best technologies for their business. Over 7,950 organizations, including 20th Century Fox, JetBlue, Nordstrom, Slack, Teach for America, and Twilio, trust Okta to help protect the identities of their workforces and customers. Position Description We are seeking an experienced Senior Software Engineer to play a key role in building and scaling the Okta Recovery Vault (ORV) . This team is responsible for Okta's enterprise-grade soft-delete and object recovery capability, designed to protect critical identity objects (Users and Groups) from accidental or malicious deletion. As a Senior Engineer, you will own the technical design, implementation, and operational reliability of critical components within our real-time, high-fidelity recovery system. You will solve complex engineering problems around identity preservation (UUIDs) and relationship restoration—including group memberships, app assignments, and password hashes—ensuring our customers can seamlessly recover from data loss events. Job Duties and Responsibilities Feature Execution: Drive the technical design and end-to-end implementation of complex features, such a

javasqlmysql
View job →
P
24 days ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect

REMOTEpythonawsmachine learning
View job →
O
24 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know Okta Okta is The World’s Identity Company. We free everyone to safely use any technology anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box we’re looking for lifelong learners and people who can make us better with their unique experiences. Join our team! We’re building a world where Identity belongs to you. About Technology Data and Intelligence at Okta At Okta, the Technology Data and Intelligence (TDI) team drives internal efficiency through secure, scalable, and innovative systems. TDI partners with teams across the company to build and support the infrastructure, automation, and enterprise applications that keep operations running smoothly. Focused on enabling productivity and aligning technology with business goals, TDI plays a vital role in both day-to-day operations and long-term strategic growth. The Staff Software Engineer Opportunity We are looking for a Staff Software Engineer to join our growing team in TDI and help scale our internal business solutions with a sharp focus on security, reliability, scalability, and intelligent automation. You will be responsible for designing and developing customization

javascriptpythonjava
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
25 days ago

About the Team The ChatGPT organization at OpenAI supports our mission by bringing advanced AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal image models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research, infrastructure, and product to build the systems that power image generation at global scale. Our team partners closely with researchers, product engineers, designers, and platform teams to bring state-of-the-art image capabilities to millions of users while continuously pushing the boundaries of what AI-powered creation can do. About the Role We are looking for an experienced Backend Engineer to join the Image Generation team and help build the systems that power image creation and editing across ChatGPT. You'll work on the core backend infrastructure that enables users to generate, edit, and iterate on visual content using cutting-edge multimodal AI models. This includes building highly scalable services, orchestration systems, APIs, storage platforms, and distributed infrastructure that support billions of image generations and editing workflows. You'll partner closely with product, research, and mobile teams to transform breakthrough AI capabilities into reliable, performant experiences used by millions around the world. In this role, you will: Design, build, and operate backend systems that power image generation and image editing experiences in ChatGPT. Develop scalable APIs, services, and infrastructure that support multimodal AI workflows. Optimize reliability, latency, throughput, and cost across large-scale distributed systems. Partner with researchers to productionize new im

REMOTEawsrestai
View job →
L
Lyft
📍 Toronto• Full-time• From C$136K/yr
25 days ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Our Infrastructure team is passionate about building software to solve problems at massive scale. We do this often, and when we believe our solution is worth sharing with the community, such as Envoy Proxy , we open source our ideas for the benefit of others. As a Infrastructure Engineer at Lyft, you will run our Production Infrastructure by monitoring system availability and take a holistic view of our platform health. You will build software and platforms to automate infrastructure platform operations and management. By measuring and monitoring our operations you will seek opportunities to optimize our systems in order to push our platform forward, anticipating our customers' needs in order to continually improve the platform. You will provide Lyft partner teams with operational support to help them build robust large scale distributed systems. About the Team Data Pipelines is at the heart of all critical data flowing through Lyft supporting hundreds of services that impact millions of drivers and passengers every day. Our team’s mission is to empower Lyft engineers to self-serve in building and maintaining data pipelines as needed to support products that deliver the world’s best transportation experience. We leverage a variety of technologies to store, stream and manage data making it available to our internal customers. Responsibilities: Maintain and analyze metrics from; operating systems; control planes; and applications to assist in fault detection and performance enhancement Design, develop and deploy tooling and systems that continually improve the reliability, scalability and efficiency of our platform Balance feature development speed and reliability with service-level objectives Operate and improve our Infrastructure using industry best practices and tools Participate in design and

pythonawsdocker
View job →
O
26 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description Okta is the leading independent provider of enterprise identity. The Okta Identity Cloud enables organizations to securely connect the right people to the right technologies at the right time. With over 6,500 pre-built integrations to applications and infrastructure providers, Okta customers can easily and securely use the best technologies for their business. Over 7,950 organizations, including 20th Century Fox, JetBlue, Nordstrom, Slack, Teach for America, and Twilio, trust Okta to help protect the identities of their workforces and customers. Position Description We are seeking an experienced Full Stack Senior Software Engineer to play a key role in building and scaling the Okta Recovery Vault (ORV). This team is responsible for Okta's enterprise-grade soft-delete and object recovery capability, designed to protect critical identity objects (Users and Groups) from accidental or malicious deletion. As a Senior Engineer, you will own the technical design, implementation, and operational reliability of critical components within our real-time, high-fidelity recovery system — spanning backend services and the admin-facing UI that customers use to review and restore their data. You will solve complex engineering problems around identity preservation (UUIDs) and relationship restoration—including group memberships, app assignments,

typescriptjavareact
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
28 days ago

About the Team OpenAI’s Client Platform Engineering (CPE) team delivers trusted devices at scale: secure by default, reliable by design, and effortless to use. We own platform capabilities across macOS, Windows, iOS, Android, and Linux, spanning endpoint posture and device trust, application delivery, onboarding, updates, telemetry, workflow orchestration, and employee-facing remediation. The team partners deeply with Security, Research, Applied, and specialized engineering groups to enable and protect OpenAI while reducing friction for the people advancing our mission. About the Role As an Engineering Manager for CPE, you will lead a team of engineers responsible for the strategy, delivery, and operation of OpenAI’s cross-platform client foundation. You will combine people leadership with strong technical judgment: setting direction, developing engineers, reviewing architecture and tradeoffs, and creating the operating mechanisms that turn ambiguous needs into durable platform outcomes. This is a high-leverage role at the intersection of security, reliability, developer velocity, and employee experience. CPE is a highly technical platform engineering organization delivering first-party services, automation, observability, and safe fleet operations. We’re looking for a leader who can guide its next chapter, scaling the team and its systems, partnering across the company, and raising the bar for secure, reliable, low-friction experiences across every supported platform. In this role, you will: Lead and develop a high-performing engineering team; hire thoughtfully, coach engineers, create clarity, and foster an inclusive, high-accountability culture that pushes perceived limits. Define and execute a multi-year client-platform strategy and roadmap across macOS, Windows, iOS, Android, and Linux, including how Codex and agents can reshape employee computing. Provide technical direction for endpoint posture, device trust, application delivery, device onboarding, updates,

REMOTEawskubernetesci/cd
View job →
O
29 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Company Description Okta is the leading independent provider of enterprise identity. The Okta Identity Cloud enables organizations to securely connect the right people to the right technologies at the right time. With over 6,500 pre-built integrations to applications and infrastructure providers, Okta customers can easily and securely use the best technologies for their business. Over 7,950 organizations, including 20th Century Fox, JetBlue, Nordstrom, Slack, Teach for America, and Twilio, trust Okta to help protect the identities of their workforces and customers. Position Description We are seeking an experienced Senior Software Engineer to play a key role in building and scaling the Okta Recovery Vault (ORV) . This team is responsible for Okta's enterprise-grade soft-delete and object recovery capability, designed to protect critical identity objects (Users and Groups) from accidental or malicious deletion. As a Senior Engineer, you will own the technical design, implementation, and operational reliability of critical components within our real-time, high-fidelity recovery system. You will solve complex engineering problems around identity preservation (UUIDs) and relationship restoration—including group memberships, app assignments, and password hashes—ensuring our customers can seamlessly recover from data loss events. Job Duties and Responsibilities Feature Execution: Drive the technical design and end-to-end implementation of complex features, such a

javasqlmysql
View job →
G
Gitlab
📍 Poland• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. About the role As a Staff Backend Engineer, you will provide technical leadership across your team and adjacent teams, solving the highest-scope and most complex problems in your area. You will lead large, cross-cutting backend initiatives, drive our modular architecture strategy, and define the standards that let teams move faster without compromising quality, security, reliability, or operability. This is a technical leadership role that combines deep backend expertise, systems judgment, product judgment, and influence across organizational boundaries. You will work with Product, Frontend, Infrastructure, Security, Data, Engine

REMOTEsqlpostgresqlkubernetes
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Top cities for Reliability Engineer

City links are canonicalized and require at least 20 current jobs.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.