Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.β Remote: This is a remote position, so youβll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team.... Global Compute runs Optimised Hosting, GoDaddy's global platform for all customer hosting products. Squad R is the engineering team responsible for operating, scaling, and continuously improving the OpenStack-based clouds that power that platform. We treat reliability as an engineering problem: we automate toil away, we plan capacity ahead of demand, and we instrument everything so that we understand our systems before they surprise us. As an SRE III on the team, you'll be a senior technical contributor who others lean on for the hard problems. What you'll get to do... Operate and scale GoDaddy's cloud infrastructure, including our OpenStack-based hosting platform. You'll troubleshoot and improve services spanning compute, networking, and storage in large-scale production environments. Drive the OpenStack migration. Help move customer hosting workloads onto the platform safely β designing and executing migration tooling, validation, and rollback strategies that protect customer experience. Work within a large-scale global hosting environment supporting thousands of servers and customer workloads across multiple regions. Eliminate toil through automation. Build and maintain automation in Python and Puppet to replace manual operational work. Treat repeated manual effort as a bug to be fixed. Strengthen observability. Improve monitoring, alerting, and dashboards so that signal reaches the right engineer at the right time, and so that we can reason about system behavior from data. Participate in on-call and incident response. T
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities Β· Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What youβll do at Jamf: At Jamf, we empower people to be their best selves and do their best work. As a Site Reliability Engineer II, youβll help us balance development velocity with the reliability our customers depend on. Youβll work with engineering teams to implement how their services are measured, investigate production issues across the stack, and turn what you learn into automation, tooling, and documentation that improves reliability. Youβll use agentic development tools as part of your everyday practice, delegating well-bounded tasks, verifying the output, and contributing to the shared context that makes AI effective for the whole team. This is a hands-on individual contributor role at the intersection of Engineering, Product, Customer Success and Technical Support, where youβll take ownership of part of your teamβs systems, deliver reliability improvements end-to-end within your team and grow toward broader influence. This role if offered as remote in Australia. You may be required to work periodically at a Jamf office or collaborative work location with other Jamf employees in your area for certain events or moments that matter. We are only able to accept applications for those based in Australia and unrestricted work rights in Australia. At this time, we are unable to offer employer sponsorship or support employer-administered work authorisation for this position. # LI -Remote What you can expect to d
Our mission and customers: We are creating the freedom for SMEs to succeed by delivering Europe's leading finance workspace with banking at its core, augmented by financial tools. We are proud to be rated 4.8 on Trustpilot, based on 55,000+ reviews. Our culture puts customer satisfaction at the core of what we do, as proven by our Net Promoter Score of 75 (more about our culture here). Our journey: Founded in 2017 by Alexandre and Steve, Qonto has grown to 1,600+ Qontoers serving over 600,000+ customers across 8 European countries. We have been profitable since 2023, and we are just getting started. Our beliefs: We hire for skills and potential. With 80+ nationalities, 45% women, of which 56% of women in our leadership team, diversity isn't a program; It's who we are. We've built a discrimination-free hiring process because the best teams are built on merit. AI at Qonto: AI is deeply embedded in how we work (here) - Every Qontoer gets unlimited access to the best AI tools. We want people who experiment without waiting for permission, push AI beyond the obvious, know when to trust it, and when to question it. ------------------------------------------------------------------------------------------------------ Join us as a Site Reliability Engineer on our Storage team to keep the databases Qonto runs on resilient, safe, and always available. You'll operate and improve our banking-grade storage infrastructure β PostgreSQL, Redis, Kafka and Elasticsearch β and help make it self-serve for backend teams, under the guidance of Damien, our Storage Engineering Manager. β‘οΈ What you'll do Operate and safeguard critical storage infrastructure: You'll run PostgreSQL, Redis, Kafka and Elasticsearch in production, keeping banking-grade data safe and available. Own incident response and root-cause analysis: You'll be the last line of defense when a hard query or performance problem needs solving. Build our banking-license compliance roadmap: You'll define and prioritize backup a
Site Reliability Engineer β . Apply via Workday.
Supply Reliability Engineer - Millennium Space Systems β USA - El Segundo, CA. Apply via Workday.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale β and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure β metrics, logging, dashboards, and alerting β as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con
Site Reliability Engineer - Comcast Technology Solutions β Great Britain - London, 1 St Giles High St. Apply via Workday.
Site Reliability Engineer, Streaming HUB - FreeWheel β VA - Reston, 11951 Freedom Dr Ste 900. Apply via Workday.
Site Reliability Engineer β CO - Centennial, 4100 E Dry Creek Rd. Apply via Workday.
Site Reliability Engineer, Streaming HUB - FreeWheel β VA - Reston, 11951 Freedom Dr Ste 900. Apply via Workday.
Synthesia is the worldβs leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on β AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge β vendor relationships, cost management, and incident response β lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence β take custody of the incident process: on-call quality, resp
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management β and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable β not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted. This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority. What You'll Own Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions Own and evolve the Operational Readiness Review (ORR) process β conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. Weβre training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. As a Site Reliability Engineer you will: Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments. Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. Take steps required to ensure we hit defined SLOs, including pa
About PostHog Product development used to mean manually writing code, running analysis, diagnosing bugs, and rolling out changes using dozens of tools. PostHog is the only platform that acts like a co-pilot for you (and your AI agents) to do it all β autonomously. We started with open-source product analytics, launched out of Y Combinator's W20 cohort . We've since shipped more than a dozen products , including: PostHog Code , the only AI devtool that understands your product, not just your codebase. A built-in data warehouse , so users can query product and customer data together using custom SQL insights. PostHog AI , an AI-powered analyst that answers product questions, helps users find useful session recordings, and writes custom SQL queries. We are: Product-led . More than 450,000 organizations have installed PostHog, mostly driven by word-of-mouth. We have intensely strong product-market fit. Default alive . Revenue is growing incredibly quickly, and we're very efficient. We raise money to push ambition and grow faster, not to keep the lights on. Well-funded. We've raised more than $180m from some of the world's top investors. We're set up for a long, ambitious journey. We're focused on building an awesome product for end users, hiring exceptional teammates, shipping fast, and being as weird as possible . Things we care about Transparency: Everyone can read about our roadmap, how we pay (or even let go of) people, our strategy, and how we work, in our public company handbook . Internally, we share revenue, notes and slides from board meetings, and fundraising plans, so everyone has the context they need to make good decisions. Autonomy: We donβt tell anyone what to do. Everyone chooses what to work on next based on what's going to have the biggest impact on our customers, and what they find interesting and motivating to work on. Engineers lead product teams and make product decisions . Teams are flexible and easy to change when needed. Shipping fast: Why not n
About PostHog Product development used to mean manually writing code, running analysis, diagnosing bugs, and rolling out changes using dozens of tools. PostHog is the only platform that acts like a co-pilot for you (and your AI agents) to do it all β autonomously. We started with open-source product analytics, launched out of Y Combinator's W20 cohort . We've since shipped more than a dozen products , including: PostHog Code , the only AI devtool that understands your product, not just your codebase. A built-in data warehouse , so users can query product and customer data together using custom SQL insights. PostHog AI , an AI-powered analyst that answers product questions, helps users find useful session recordings, and writes custom SQL queries. We are: Product-led . More than 450,000 organizations have installed PostHog, mostly driven by word-of-mouth. We have intensely strong product-market fit. Default alive . Revenue is growing incredibly quickly, and we're very efficient. We raise money to push ambition and grow faster, not to keep the lights on. Well-funded. We've raised more than $180m from some of the world's top investors. We're set up for a long, ambitious journey. We're focused on building an awesome product for end users, hiring exceptional teammates, shipping fast, and being as weird as possible . Things we care about Transparency: Everyone can read about our roadmap, how we pay (or even let go of) people, our strategy, and how we work, in our public company handbook . Internally, we share revenue, notes and slides from board meetings, and fundraising plans, so everyone has the context they need to make good decisions. Autonomy: We donβt tell anyone what to do. Everyone chooses what to work on next based on what's going to have the biggest impact on our customers, and what they find interesting and motivating to work on. Engineers lead product teams and make product decisions . Teams are flexible and easy to change when needed. Shipping fast: Why not n
Get new reliability engineer jobs by email
Daily job updates Β· Unsubscribe anytime