Location: South Jordan, UT (This role is on-site) Salary: $56,000 - $58,000 USD Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. As part of the mthree Alumni program, mthree has an exciting and exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, e
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. We’re hiring a Manufacturing Reliability Engineer to own production test for our robots at our contract manufacturer: you’ll design and run robust end-to-end test protocols, provision fleets of robots for production, and own the KPIs that define production quality. This role is based in Austin, TX. However, the position will require 50% travel to the Milwaukee, WI area and requires close collaboration across software, hardware, operations, and product engineering teams. Key Responsibilities End-to-end test process ownership. Create, validate, and maintain production test protocols and gating criteria from incoming inspection through final test and shipment. Provisioning of bots. Design and operate provisioning flows (imaging, firmware deployment, configuration, validation) and the tooling/fixtures needed to provision and handoff robots for production. KPIs and continuous improvement. Own key production metrics — First Pass Yield (FPY), cycle time, and test coverage — and drive continuous improvements to meet throughput and quality targets. Test automation & infrastructure. Architect, implement, and maintain automated test frameworks, harnesses, and test rigs used at the CM site. Ensure tests are stable, fast, and provide actionable failure data. Cross-functional escalation & RCA. Lead root-cause analysis for field and production failures; coordinate corrective actions with design, firmware, and CM engineering to close quality loops. On-site production leadership. Be the onsite technical authority at the contract manufacturer: train operators, debug failures on the line, and continuously refine processes with CM partners. What Success Looks Like Improved FPY and reduced rework rates across production builds. Reduced per
We believe communication belongs to everyone. We exist to democratize phone service. TextNow is evolving the way the world connects and that's because we're made up of people with curious minds who bring an optimistic, yet critical lens into the work we do. We're the largest provider of free phone service in the nation. And we're just getting started. Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between! This role is about impact at scale. You’ll shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded into everyday engineering practice. Using AI is not optional, it’s expected. From design and architecture to implementation, testing, debugging, documentation, and operational analysis, you will actively leverage AI tools to increase velocity, improve code quality, and make better technical decisions. We provide a robust suite of AI-powered development tools and workflows to support you, and we expect you to continuously evolve how you use them to raise the bar for efficiency, clarity, and product excellence across the organization. What You'll Do Ensure System Reliability: Design, build, and maintain scalable, resilient, and highly available systems to support TextNow’s infrastructure and services. Automation & Infrastructure as Code: Develop and maintain automation using Terraform, Ansible, and other tools to enable efficient deployment, scaling, and operations of cloud-based systems (AWS preferred). Incident Response & On-Call Support: Participate in an on-call rotation, troubleshoot issues, and drive incident resolution to minimize downtime and improve syste
Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida/ Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite
Title: Senior Site Reliability Engineer - I, Product Area Focus Location: Noida (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering teams. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedit
Title: Staff Site Reliability Engineer, Product Area Focus Location: Noida / Bangalore (Hybrid) Summary of role Own availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work alongside your global SRE team, executing on projects in your product-area specific reliability roadmap, to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture, and increase feature velocity of our developers Work closely with multiple teams to optimize the operations of their microservices - and improve the lives of the engineers within your product area engineering team. Responsibilities Support the engineering teams within your product area by maintaining and executing a reliability roadmap of opportunities for improvement for reliability, maintainability, security, efficiency, and velocity - and help for realizing those opportunities. Collaborate with development infrastructure, Global SRE, and your product area engineering teams to establish and continually refine your reliability roadmap. Participate in defining, evolving, and managing SLOs for several teams within your product area. Participate in on-call rotations within your product area to understand operations workload so you can continually work to improve the on-call experience and reduce operational workload for running microservices and related components. Complete projects to optimize and tune on-call experience for your engineering teams. Continually improve the lifecycle of microservices and architectural components from inception and design, through deployment, operation, and refinement. Write code and automation to reduce operational workload, increase efficiency, improve security posture, eliminate toil, and enable Sumo’s developers to deliver features more rapidly. Work closely with the developer infrastructure teams to expedite
Role: Application Reliability Engineer Location: Gurgaon Who we are Graviton Research Capital is a privately funded quantitative trading firm striving for excellence in financial markets research. We trade across a multitude of asset classes and trading venues using a diverse range of concepts, from time series analysis and stochastic models to machine learning and statistical inference. We analyse terabytes of data to identify pricing anomalies and drive innovation in financial markets. Key Responsibilities and Deliverables The ideal candidate will possess a strong background in technical support, with a passion for problem-solving and a commitment to excellence. As an Application Reliability Engineer, you will be responsible for: Monitor production services and respond quickly to alerts, incidents, and outages to ensure smooth operation and minimal downtime. Monitor trading systems and infrastructure., Triage issues across trading support services, databases, and infra; escalate and coordinate with the right owners, and drive root-cause analysis and ensure fixes are implemented for long-term stability. Serve as the first line of defense for trading operations. Proactively identify, address recurring issues, and build automation to reduce manual intervention. Improve observability by enhancing monitoring, logging, and alerting systems. Develop and maintain operational runbooks and SLO/SLA metrics. Eligibility and Required Skills Possess a degree in a highly analytical field, such as Engineering, or Computer Science 2-5 years of experience in Python, Shell/Bash scripting. Experience with Linux and shell/bash online tools. Hands-on experience with databases (SQL, NoSQL) Strong problem-solving and analytical skills Excellent communication skills Ability to remain calm and analytical under production pressure Good to have: Familiarity with monitoring/alerting stacks (Prometheus, Grafana, ELK, etc.) Familiarity with distributed messaging (Kafka) and caching systems (Red
JOB TITLE Site Reliability Engineer A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open-source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO You will play a highly critical operational role where you will apply a combination of software and systems engineering skills to develop and maintain a complex set of distributed, real-time systems that serve critical stakeholders in Point72’s Global Macro business. You will focus on optimizing the operations of existing systems and infrastructure in an efficient manner, through a strict adherence to automation and tooling Specifically, you will: Build out foundational technical components of an extensive SRE program across multiple complex systems, both new and existing • Collaborate with our development and quant teams to ensure that ongoing change is consistent with a pre-determined, measurable set of SLOs spanning multiple complex user interactions with our systems • Monitor system capacity and performance, identifying and addressing potential future bottlenecks and sources of instability before they become impactful to our stakeholders • Review and provide feedback on automation code developed by peers to maintain high standards of code quality and efficiency • Troubleshoot and resolve system issues, analyzing their impact on infrastructure and service operations • Participate in or lead design reviews with peers and stakeholders, evaluating and selecting the best technologies and automation strategies for our needs WHAT’S REQUIRED We are looking for highly motivated, proactive engineers
A Career with Cubist Cubist Systematic Strategies, an affiliate of Point72, deploys systematic, computer-driven trading strategies across multiple liquid asset classes, including equities, futures, and foreign exchange. The core of our effort is rigorous research into a wide range of market anomalies, fueled by our unparalleled access to a wide range of publicly available data sources. What You’ll Do We are passionate about data. We collaborate to build elegant, effective, scalable, and highly reliable solutions to empower predictive modeling in finance. You will join a team that plays a vital role in ensuring the smooth day-to-day implementation of a large research infrastructure and the timely delivery of comprehensive and error-free data to Cubist’s portfolio managers across the globe. Specifically, you will: Serve as a frontline owner for thousands of mission-critical data ETL pipelines that power trading and investment decision-making, ensuring reliability, accuracy, and timeliness. Actively manage and resolve data incidents in a fast-paced trading environment, partnering closely with portfolio managers, data scientists, and external data vendors. Design and build tooling, automation, and robust documentation to improve operational efficiency, scalability, and data quality across the platform. Play a hands-on role in daily data operations, including data validation, remediation, and enrichment, with opportunities to continuously improve and modernize workflows through engineering best practices. What’s Required Bachelor’s degree with a focus in computer science or a related field. Strong proficiency in SQL Server and Python programming, with experience in AWS and both Windows and Linux environments; familiarity with Databricks is a plus. Exceptional attention to detail with a strong appreciation for well-defined processes and systems. 3+ years of experience in a client-facing support or operations role. Excellent organizational, communication, and interpersonal
JOB TITLE Data Reliability Engineer A CAREER WITH CUBIST Cubist Systematic Strategies, an affiliate of Point72, deploys systematic, computer-driven trading strategies across multiple liquid asset classes, including equities, futures, and foreign exchange. The core of our effort is rigorous research into a wide range of market anomalies, fueled by our unparalleled access to a wide range of publicly available data sources. What you’ll do Ensure smooth day-to-day implementation of a large research infrastructure and the timely delivery of comprehensive and error-free data to Cubist’s portfolio managers across the globe Serve as a frontline owner for mission-critical data ETL pipelines that power trading and investment decision-making, ensuring reliability, accuracy, and timeliness. Actively manage and resolve data incidents in a fast-paced trading environment, partnering closely with investment professionals, data scientists, and external data vendors. Design and build tooling, automation, and robust documentation to improve operational efficiency, scalability, and data quality across the platform. Play a hands-on role in daily data operations, including data validation, remediation, and enrichment, with opportunities to continuously improve and modernize workflows through engineering best practices. What’s REQUIRED Bachelor’s degree in computer science or a related field. Strong proficiency in SQL Server and Python programming, with experience in AWS and both Windows and Linux environments. Exceptional attention to detail with a strong appreciation for well-defined processes and systems. 3+ years of experience in a client-facing support or operations role. Excellent organizational, communication, and interpersonal skills. Commitment to the highest ethical standards About point72 Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Poin
About the Team DoorDash Labs, established in 2018, serves as the innovation hub for DoorDash, focusing on developing automation and robotics solutions to enhance last-mile logistics. The team's mission is to create technologies that support and augment human networks, aiming to improve efficiency for Dashers, merchants, and consumers alike. We’re ruthlessly focused on business impact. We are a highly senior team composed of former pioneers from a variety of different robotics industries. As of 2025, DoorDash has completed 10B lifetime deliveries. We’re focused on how to do the next 10B even better. About the Role We are seeking a highly motivated Senior Reliability & Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the system and component levels. You will partner closely with EE, ME, and Autonomy teams to translate mission needs into robust, reliable hardware. The ideal candidate thrives in a fast-moving, cross-functional environment where reliability and test rigor determine program success. You will be hands-on in developing test methods and equipment to uncover failures before they happen in the field. You will partner closely with EE, ME, and Autonomy teams to translate mission needs into robust, reliable hardware. The ideal candidate thrives in a fast-moving, cross-functional environment where reliability and test rigor determine program success. You’re excited about this opportunity because you will… Architect and implement rigorous validation strategies, utilizing Python scripts for automation while leveraging CAD and shop tools to engineer bespoke test fixtures and hardware rigs. Oversee experimental execution across internal facilities and external laboratories, maintaining technical mastery over vibration tables, environmental chambers, DAQ systems, and ingress protection testing. Translate high-level vehicle reliability requirements into granula
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta’s TDI Network Engineering team is responsible for the global corporate network, building and supporting a high-performing, reliable network at scale. As a member of this team, you will have a direct impact on network design, deployment, and reliability, enabling our employees to work effectively from any location globally. Your role ensures the overall security and integrity of our corporate network by leveraging network security best practices, innovative products, and rigorous security validation. Reporting to the Network Engineering Manager, this operations-focused role is distinct from core Network Engineering and Network Security, centering primarily on operational execution—including responding to alerts, maintaining service availability, and ensuring system health across our global enterprise network. You will drive the strategic reduction of systemic toil and technical debt across multiple teams, applying a systems-level perspective and leveraging deep expertise in Distributed Systems, Networking fundamentals, Infrastructure as Code, and observability to architect scalable platforms and lead technical efforts to ensure an "Always Secure. Always On." environment. You will own multi-quarter objectives and establish long-term strategies for network reliability. What you'll be doing : Design and Own the resilience, health and availability of our entire global corporate network domain, managing operational responsibilities such as responding to aler
Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team The Linux Database Services team is responsible for operating GoDaddy's on-premise database infrastructure, supporting Percona MySQL, PostgreSQL, and Cassandra. We provide database services for internal applications and support much of GoDaddy's hosting infrastructure. Our scale is significant: thousands of database instances hosting hundreds of thousands to millions of individual databases! We are a globally distributed team with members across the EU, North America, South America, and South Asia (India), and we are growing our presence in Pune. As a Site Reliability Engineer, you will play a key role in keeping our database fleet healthy, performant, and observable. You will work across Percona MySQL, PostgreSQL, and Cassandra environments, handling both day-to-day operational work and longer-term engineering improvements. This role is ideal for someone who is passionate about building reliable, scalable infrastructure, enjoys solving complex operational challenges, and thrives in an environment where automation and continuous improvement are core to how work gets done. What you'll get to do... Maintain and improve the health of database servers and instances across a large-scale fleet Build and enhance telemetry and observability for databases and hosts Identify and implement performance improvements across database systems Execute an
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: As a Site Reliability Engineer, you will play a crucial role in ensuring the smooth operation of all user-facing services and other Anyscale production systems. Anyscale values diversity and inclusion, and we encourage applications from individuals of all backgrounds. This includes processes for provisioning, negotiating prices, managing costs, seeing opportunities for teams to reduce wastage by finding applications across the company. You will apply sound engineering principles, operational discipline, and mature automation to our environments and the Anyscale codebase as we scale. As part of this role, you will: Develop a unified perspective on how cloud components are utilized across the company, taking into account diverse needs and requirements. Ensure that deployment methodologies align with the company's reliability goals. Build systems that promote understanding of production environments, facilitating quick identification of issues through robust observability infrastructure for metrics, logging, and tracing. Create monitoring and alerting systems at different levels, enabling teams to easily contribute and enhance the overall monitoring capabilities. Establish testing infrastructure to s
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime