Jobiba hiring network

Reliability Engineer Jobs

2,049 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations Austin, US Responsibilities The Argo team was formed to own a very important aspect of Cloudflare's systems: enable more reliable network connectivity for Cloudflare’s products than the Internet itself provides. Almost all products in Cloudflare’s portfolio are or will be powered by Argo technology, including CDN, Spectrum, Magic Transit, Stream, Workers, Workers AI, R2, WARP, and more. As a member of the Argo team, you’ll be a

sqlpostgresqlaws
View job →
A
1mo ago

We're looking for an Engineering Manager who combines strong technical judgment with people leadership to help build the systems and team practices that keep Asana resilient at scale. This role is a great fit for someone who enjoys turning broad reliability problems into clear priorities, and can balance foundational engineering work with iterative delivery. You'll help shape what Platform Reliability means at Asana while building a team that ships durable systems, strong operational practices, and high-trust partnerships. You will define the reliability roadmap for a rapidly growing global platform, transforming reliability into a core architectural advantage. You'll partner closely with platform engineering, infrastructure, and product teams in Warsaw, Reykjavik and San Francisco to protect Asana under real-world load, improve how traffic and failure modes are handled, and ensure reliability is designed in rather than added later. This is a role for someone who can lead through influence, coach engineers, and raise the bar for execution and collaboration. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you'll achieve Build and lead a new Platform Reliability team, hiring and developing engineers while setting a clear standard for collaboration, ownership, technical excellence and growth. Partner with technical leaders to define the roadmap for core reliability systems such as load shedding, rate limiting, circuit breakers, traffic controls, and other platform guardrails. Establish and evangelize best-in-class operating practices for incident response, postmortems, and proactive risk management through SLOs a

awskubernetesrest
View job →
A
Asana
📍 Warsaw• Full-time• $372K – $432K/yr
1mo ago

Asana’s rapid growth brings new challenges in keeping our systems fast, reliable, and resilient. As our product evolves, we’re making a major investment in reliability – and building a brand new SRE team in Warsaw is a key part of that strategy. This is your chance to help shape it from day one. This isn’t a traditional “ops” role – we’re looking for strong software engineers who are passionate about building reliable, distributed systems. You’ll work closely with a small SRE team in San Francisco, infrastructure engineers in Reykjavik, and an established infrastructure team in Warsaw. Warsaw will be a significant hub for our future infrastructure engineering and operations. As one of the first engineers here, you’ll have a real say in how we build reliable infrastructure, manage incidents, and support the rest of the company. This role is based in our Warsaw office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do, and your recruiter can share more about the in-office requirements. We offer a Contract of Employment (UoP) for our employees in Poland. What you’ll achieve Influence the future of Asana’s SRE practice, especially as we grow the Warsaw team. Lead reliability-focused projects across our stack – from infrastructure to tooling to incident response. Define and implement Asana’s incident management process – we’re investing here, and you’ll help shape how it works. Build internal platforms and frameworks that help other teams improve the reliability of their services. Be part of (and help shape) a sustainable on-call rotation – shared across teams in Warsaw, San Francisco, and Reykjavik. On average, we handle ~1 page per day, but it’s not constant, and we care about keeping things sane. Work with our stack: AWS, Kubernetes (EKS), Datadog, MySQL (RDS), ElasticSearch (OpenSearch), Redis

typescriptpythonsql
View job →
O
1mo ago

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role As OpenAI continues to grow, we are looking for experienced, problem-solving engineers to ensure our systems scale. Our success depends on our ability to quickly iterate on products while also ensuring that they are performant and reliable. You will work in a deeply iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. Successful candidates will play a crucial role in ensuring the reliability, scalability, and performance of our systems as we continue to expand. As a reliability expert, you will be at the forefront of maintaining and enhancing the stability, scalability, and performance of our rapidly evolving infrastructure. You will work closely with cross-functional teams, including software engineers, product managers, and data scientists, to build and maintain resilient systems that can handle our growing user base and workload. In this role, you will: Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands. Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable. Build and maintain automation tools to streamline repetitive tasks and improve system reliability. Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources. Implement fault-tolerant and resilient

awskubernetesrest
View job →
🔔

Get new reliability engineer jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More reliability engineer opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.

Countries hiring Reliability Engineer

Country links use the same curated canonical inventory as Jobiba sitemaps.