Jobiba hiring network

Reliability Engineer Jobs

2,028 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time• Remote
24 days ago

About the Team The ChatGPT organization at OpenAI supports our mission by building products that bring cutting-edge AI capabilities to hundreds of millions of users worldwide. The Image Generation team is responsible for one of the fastest-growing experiences in ChatGPT, enabling users to create, edit, and transform images through natural language. Recent advances in our multimodal models have dramatically improved image quality, instruction following, editing precision, consistency, and text rendering, unlocking entirely new creative and professional workflows. We work at the intersection of research and product, partnering closely with researchers, designers, product managers, and platform engineers to bring state-of-the-art image generation capabilities to life across ChatGPT and our mobile applications. Millions of users rely on these experiences every day to create, communicate, learn, and build. About the Role We are seeking an experienced Android Software Engineer to build and improve image generation experiences within the ChatGPT Android app. You will help define how users create, edit, and interact with visual content powered by the latest multimodal AI models. This is an opportunity to work on a highly visible product area, translating cutting-edge AI capabilities into intuitive, performant, and delightful mobile experiences used by millions around the world. ChatGPT's Android app already enables users to generate and transform images directly from their devices, and we're just getting started. In this role, you will: Build and ship new Android features that power image generation and image editing experiences. Create intuitive user experiences that make advanced AI capabilities feel seamless and accessible. Collaborate closely with Product, Design, Research, and Engineering teams to bring new multimodal capabilities to production. Drive improvements in app performance, reliability, architecture, testing, and developer tooling. Optimize media-heavy workfl

REMOTEawsrestai
View job →
R
24 days ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world’s biggest financial problems. We’re looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn’t a place for complacency, it’s where ambitious people do the best work of their careers. We’re a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The International engineering team's mission is to expand Robinhood's products globally and scale our platform to support millions of new customers in diverse markets. We build and maintain experiences across onboarding, funding, account setup, localized product experiences, growth incentives, and platform capabilities to operate in new jurisdictions. Our engineering team partners closely with product, design, compliance, marketing teams, along with an array of platform engineering teams to launch new products and optimize user acquisition globally. We prioritize technical rigor, system reliability, and quick execution to deliver reliable products to our growing customer base. Our culture is centered on clear communication, strong team partnership, and a commitment to helping people manage their financial lives! As an Engineering Manager , you will lead a team of software developers to drive some of Robinhood’s most important international launch initiatives. You will be responsible for driving the architecture and execution required to launch crypto and other financial products in new regions, while shaping the reusable systems that make future launches faster, safer, and m

O
OpenAI
📍 San Francisco• Full-time• Remote
26 days ago

About the Team OpenAI develops models that can reason through complex problems and hardware designed for the demands of advanced AI. AI for Chips connects these efforts: applying increasingly capable AI systems to the work of semiconductor engineering. Our goal is to help engineers develop better chips and shorten design cycles. This work brings research, model training, and hardware expertise together to build tools that engineers can use on real designs, with correctness and measurable performance at the center. About the Role We’re hiring a Software Engineer to build the research infrastructure and tooling that help OpenAI models design silicon. You’ll turn chip-design workflows into reliable environments for reinforcement learning and evaluation, and make it easier for researchers to run experiments and iterate on new ideas. You’ll move between software engineering, tool integration, and open research problems. We value strong coding fundamentals, clear technical judgment, and independent execution. Prior chip-design experience is helpful, but you can learn the domain alongside the team’s hardware specialists. In this role, you will: Build and maintain infrastructure for reinforcement learning environments, evaluations, and long-running experiments. Integrate electronic design automation (EDA) tools into workflows for RTL generation, verification, and physical design optimization. Improve experiment reliability, reproducibility, observability, and performance; debug failures across tools, services, and infrastructure. Develop tooling and model harnesses that let researchers test ideas quickly and measure correctness and power, performance, and area (PPA). Collaborate with researchers and engineers to turn successful experiments into reusable systems and training workflows. Own ambiguous projects end to end, communicate progress, and use results to guide the next iteration. You might thrive in this role if you: Have strong software engineering fundamentals, with

REMOTEpythonawsrest
View job →
O
26 days ago

About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol

awsrestai
View job →
O
26 days ago

About the Team The ChatGPT Library team is building the place where people can save, organize, rediscover, and build on the content they create with ChatGPT. Our goal is to make ChatGPT more useful over time by helping users seamlessly return to important files, images, conversations, and other content across their devices. The team works at the intersection of product engineering, design, and AI research to create intuitive, personalized experiences that make users’ content easy to find and act on. On Android, we are focused on delivering fast, reliable, and deeply native experiences that put a user’s evolving body of work at their fingertips. About the Role We are looking for a senior Android engineer to help build the future of ChatGPT Library on mobile. You will own high-impact product experiences across the Android stack, shaping how millions of people save, organize, discover, and interact with their content in ChatGPT. In this role, you will: Build and ship new Android experiences that help users easily access, organize, and build on the content they create with ChatGPT. Own features end to end—from early product exploration and technical design through implementation, experimentation, launch, and iteration. Develop scalable, maintainable foundations that allow Libraries experiences to evolve quickly as new AI capabilities emerge. Improve the architecture, performance, reliability, and responsiveness of content-rich experiences across a wide range of Android devices. Thoughtfully integrate Android platform capabilities to create experiences that feel intuitive and native to mobile. Partner closely with product, design, research, data science, and engineering teams to translate emerging AI capabilities into useful, polished products. Help establish technical direction and raise the quality bar for Android development across the team. How We Work We care deeply about building products that are intuitive, useful, and trustworthy. We move quickly from ideas to wo

redisawsrest
View job →
D
Datadog
📍 Boston• Full-time• From $100K/yr
27 days ago

We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data structures,

kubernetesrestai
View job →
D
Datadog
📍 Madrid• Full-time
27 days ago

We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact for customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with guidance from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Expected to graduate in 2027 with a degree in Computer Science, Software Engineering, or a related technical field from a university in Spain Demonstrate strong computer science fundamentals, including data structures, algorithms, and software

kubernetesrestai
View job →
T
Twilio
📍 India• Full-time• Remote
28 days ago

Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as our next Senior Engineering Manager , Twilio’s Segment team. About the job As a Senior Engineering Manager on the Twilio Segment Data platform/ pipelines team, you’ll build and scale systems that process several hundred thousands of data points per second. You will lead the development of high-scale ingestion and data processing systems You'll guide the team in designing, operating and maintaining complex distributed systems, ensuring reliability, performance, and cost-efficiency while querying petabytes of data for our customer data platform (CDP). Responsibilities In this role, you’ll: Own and deliver robust, high-scale routing experiences for Data platforms & pipelines for Twilio Segment. Champion team growth and success, prioritizing mentorship and individual development. Architect and operate always-available, complex distributed systems in cloud environments. Guide technical decisions, articulating trade-offs between cost, p

REMOTEawsdockerkubernetes
View job →
W
Writer
📍 New York City• Full-time• Remote
1mo ago

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need. This is a hybrid position, based out of our New York City, San Francisco, Seattle, or London hubs. You'll report to our director of engineering. 🦸🏻‍♀️

REMOTEpythonawsazure
View job →
O
OpenAI
📍 San Francisco• Full-time• Remote
1mo ago

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

REMOTEawsrestai
View job →
W
Writer
📍 London• Full-time• Remote
1mo ago

🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need. This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering. 🦸🏻‍♀️ What you'll do Technical

REMOTEpythonawsazure
View job →
S
Stripe
📍 New York• Full-time• $190.4K – $285.6K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Lead the technical design and architecture of major platform initiatives, author design documents and build consensus across engineering teams. Define technical roadmaps for complex, multi-quarter projects that span multiple teams. Make critical architectural decisions for company documentation infrastructure, balancing scalability, reliability, and developer experience. Evaluate and set direction for integrating emerging technologies, including AI/LLM capabilities, into company documentation platforms and authoring tools. Establish and evolve engineering standards, best practices and technical guidelines for the team and broader organization. Partner with engineering teams across the company to understand documentation needs and design integrated solutions. Design, build and maintain scalable, reliable and performant services and systems. Contribute high-quality code across the full stack and navigate codebases with different languages and tools. Debug and resolve complex production issues and improve system reliability. Take ownership of system health and incident response. Who you are Minimum requirements Must have a Bachelor's degree or foreign equivalent in Computer Science, Software Engineering, Engineering, or a related field, plus four (4) years of experience in Software Engineering. Must have four (4) years of experience in each of the following: - Working in a full stack environment with a foc

typescriptjavamongodb
View job →
G
Gitlab
📍 Bengaluru• Full-time
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Senior Backend Engineer About the role As a Senior Backend Engineer, you will design, implement, and evolve product capabilities while solving high-scope backend problems and influencing the technical and product direction of our teams. You will move beyond "assigned work" to actively improve the quality, reliability, and performance of our systems. You will work across product, frontend, infrastructure, data, and security boundaries, making sound architectural trade-offs, communicating complex ideas clearly in an asynchronous environment, and helping define the standards for a high-scale, global product. Why you’ll love this rol

sqlpostgresqlkubernetes
View job →
C
Coinbase
📍 Singapore• Full-time• Remote• From S$143.7K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Software Engineer, Data Layer As a Software Engineer on the Data Layer team within the Platform group, you'll shape the API (GraphQL) platform that connects every client application to Coinbase's backend services, handling the majority of user traffic across Consumer, Base, and Institutional products. You'll own the critical systems that route and serve API requests at scale, improving reliability, performance, and developer experience for hundreds of engineers building on this foundation layer. What you'll do: Own and deliver projects end to end, from scoping and system design through implementation, rollout, and production validation, driving measurable progress on the team's highest-priority initiatives. Design and build high-reliability, low-latency systems serving millions of users, tackling challenges like caching, upstream service optimization, and efficient connection management at scale. Build and improve the API framework and tooling that hundreds of engineers depend on, making it fast and easy for teams across the company to build, test, and ship independently. Drive operational excellence across T0 services: own SLOs, improve observability, lead incident response, and reduce operational toil so the team can invest in high-leverage work. Partner cross-functionally with product engineering teams to design API schemas, support service launches, and ens

REMOTEjavaawsgraphql
View job →
C
Coder
📍 United States• Full-time• Remote
1mo ago

As an Engineering Manager on Coder’s Core Workspaces team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction while growing the team and keeping execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Workspaces organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with React and TypeScript . Experience with Go . Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS . Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP , agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building abstractions across multiple model providers. Deep experience with AWS, Kube

REMOTEtypescriptreactaws
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.