Jobiba hiring network

Data Reliability Engineer Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current data reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Point72
📍 SingaporeFull-time
3 days ago

A Career with Cubist Cubist Systematic Strategies, an affiliate of Point72, deploys systematic, computer-driven trading strategies across multiple liquid asset classes, including equities, futures, and foreign exchange. The core of our effort is rigorous research into a wide range of market anomalies, fueled by our unparalleled access to a wide range of publicly available data sources. What You’ll Do We are passionate about data. We collaborate to build elegant, effective, scalable, and highly reliable solutions to empower predictive modeling in finance. You will join a team that plays a vital role in ensuring the smooth day-to-day implementation of a large research infrastructure and the timely delivery of comprehensive and error-free data to Cubist’s portfolio managers across the globe. Specifically, you will: Serve as a frontline owner for thousands of mission-critical data ETL pipelines that power trading and investment decision-making, ensuring reliability, accuracy, and timeliness. Actively manage and resolve data incidents in a fast-paced trading environment, partnering closely with portfolio managers, data scientists, and external data vendors. Design and build tooling, automation, and robust documentation to improve operational efficiency, scalability, and data quality across the platform. Play a hands-on role in daily data operations, including data validation, remediation, and enrichment, with opportunities to continuously improve and modernize workflows through engineering best practices. What’s Required Bachelor’s degree with a focus in computer science or a related field. Strong proficiency in SQL Server and Python programming, with experience in AWS and both Windows and Linux environments; familiarity with Databricks is a plus. Exceptional attention to detail with a strong appreciation for well-defined processes and systems. 3+ years of experience in a client-facing support or operations role. Excellent organizational, communication, and interpersonal

pythonsqlaws
View job →
P
Point72
📍 BengaluruFull-time
3 days ago

JOB TITLE Data Reliability Engineer A CAREER WITH CUBIST Cubist Systematic Strategies, an affiliate of Point72, deploys systematic, computer-driven trading strategies across multiple liquid asset classes, including equities, futures, and foreign exchange. The core of our effort is rigorous research into a wide range of market anomalies, fueled by our unparalleled access to a wide range of publicly available data sources. What you’ll do Ensure smooth day-to-day implementation of a large research infrastructure and the timely delivery of comprehensive and error-free data to Cubist’s portfolio managers across the globe Serve as a frontline owner for mission-critical data ETL pipelines that power trading and investment decision-making, ensuring reliability, accuracy, and timeliness. Actively manage and resolve data incidents in a fast-paced trading environment, partnering closely with investment professionals, data scientists, and external data vendors. Design and build tooling, automation, and robust documentation to improve operational efficiency, scalability, and data quality across the platform. Play a hands-on role in daily data operations, including data validation, remediation, and enrichment, with opportunities to continuously improve and modernize workflows through engineering best practices. What’s REQUIRED Bachelor’s degree in computer science or a related field. Strong proficiency in SQL Server and Python programming, with experience in AWS and both Windows and Linux environments. Exceptional attention to detail with a strong appreciation for well-defined processes and systems. 3+ years of experience in a client-facing support or operations role. Excellent organizational, communication, and interpersonal skills. Commitment to the highest ethical standards About point72 Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Poin

pythonsqlaws
View job →
TI
THG Ingenuity
📍 Manchester, United KingdomFull-time
3 days ago

About THG Ingenuity THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer. Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Database Platform Manager Company: THG Ingenuity Location: Manchester (Head Office) Reports to: Director of Data Role Overview We are looking for a strong technical leader to run our database platform. Our estate runs on Google Cloud Platform following a recent migration to self-managed services, and there is a real opportunity here to shape the next phase. How we consolidate and modernise the estate, where managed and cloud-native services earn their place, which new technologies are worth adopting, and how the platform scales with the business are all live questions. You will lead the thinking on them and work with stakeholders across the business to agree and deliver the roadmap. You will lead our DBA team, who own every database across the group regardless of the application running on it and who run the platform as a 24/7 service. You will work hand in hand with our Data Reliability Engineering team, whose Principal Engineer is your peer and whose focus is automation, fleet reliability and SLOs. This role is weighted toward hands-on technical depth. You should be as credible in a design review or an incident bridge as you are in a planning session with senior stakeholders. Key Responsibilities Technical direction The technical roadmap for the estate. We want someone genuinely interested in emerging database technologies who evaluates them on merit and can articulate the pros and c

pythonsqlpostgresql
View job →
G
3 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

aigoexcel
View job →
O
Okta
📍 TorontoFull-timeFrom C$160K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi

javaawskubernetes
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. * This position requires the ability to access U.S. National Security information. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15 ) upon hire. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security &amp

kubernetesmachine learningartificial intelligence
View job →
Q
1mo ago

Our mission and customers: We are creating the freedom for SMEs to succeed by delivering Europe's leading finance workspace with banking at its core, augmented by financial tools. We are proud to be rated 4.8 on Trustpilot, based on 55,000+ reviews. Our culture puts customer satisfaction at the core of what we do, as proven by our Net Promoter Score of 75 (more about our culture here). Our journey: Founded in 2017 by Alexandre and Steve, Qonto has grown to 1,600+ Qontoers serving over 600,000+ customers across 8 European countries. We have been profitable since 2023, and we are just getting started. Our beliefs: We hire for skills and potential. With 80+ nationalities, 45% women, of which 56% of women in our leadership team, diversity isn't a program; It's who we are. We've built a discrimination-free hiring process because the best teams are built on merit. AI at Qonto: AI is deeply embedded in how we work (here) - Every Qontoer gets unlimited access to the best AI tools. We want people who experiment without waiting for permission, push AI beyond the obvious, know when to trust it, and when to question it. ------------------------------------------------------------------------------------------------------ Join us as a Site Reliability Engineer on our Storage team to keep the databases Qonto runs on resilient, safe, and always available. You'll operate and improve our banking-grade storage infrastructure — PostgreSQL, Redis, Kafka and Elasticsearch — and help make it self-serve for backend teams, under the guidance of Damien, our Storage Engineering Manager. ➡️ What you'll do Operate and safeguard critical storage infrastructure: You'll run PostgreSQL, Redis, Kafka and Elasticsearch in production, keeping banking-grade data safe and available. Own incident response and root-cause analysis: You'll be the last line of defense when a hard query or performance problem needs solving. Build our banking-license compliance roadmap: You'll define and prioritize backup a

postgresqlredisai
View job →
CV
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Product Reliability Engineers (PREs) are responsible for the health, performance, and stability of the services that power services at Palantir. PREs take ownership over the entire end-to-end cycle of service reliability, from responding to outages to improving codebases and building lasting solutions. You will tackle critical issues for key customers, introduce observability into complex systems, address tech debt in essential codebases, and inform strategic investments in core products. We are looking for engineers who enjoy deep-dive troubleshooting, feel strong ownership over the problems they encounter, and recognize the urgency of customer-facing outages. PREs spend the majority of their time on forward-looking product work, including but not limited to, infrastructure migrations, product contributions to improve stability and observability, and codebase enhancements that increase resilience. During periodic on-call shifts, we respond to automated alerts, investigate issues reported by customers, and share technical expertise with adjacent product teams. Whatever the technical issue or question about your service is, you'll play a central and critical role in resolving it, seeking not just a one-time fix, but a permanent solution. We provide new team members with an experienced mentor and a clear onboarding framework to set them up for success in the role.

CV
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, across both cloud & on-prem environments. Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll be the experts for the environments that you operate infrastructure in, helping partner teams build & configure their software to operate reliably within. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.

CV
Company via Lever
📍 WashingtonFull-timeHybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Product Reliability Engineers (PREs) are responsible for the health, performance, and stability of the services that power services at Palantir. PREs take ownership over the entire end-to-end cycle of service reliability, from responding to outages to improving codebases and building lasting solutions. You will tackle critical issues for key customers, introduce observability into complex systems, address tech debt in essential codebases, and inform strategic investments in core products. We are looking for engineers who enjoy deep-dive troubleshooting, feel strong ownership over the problems they encounter, and recognize the urgency of customer-facing outages. PREs spend the majority of their time on forward-looking product work, including but not limited to, infrastructure migrations, product contributions to improve stability and observability, and codebase enhancements that increase resilience. During periodic on-call shifts, we respond to automated alerts, investigate issues reported by customers, and share technical expertise with adjacent product teams. Whatever the technical issue or question about your service is, you'll play a central and critical role in resolving it, seeking not just a one-time fix, but a permanent solution. We provide new team members with an experienced mentor and a clear onboarding framework to set them up for success in the role.

CV
Company via Lever
📍 United KingdomFull-timeHybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Forward Deployed Reliability Engineer (FDRE), you ensure the stability and reliability of mission-critical workflows built on Palantir software. You gather signal by going on call — resolving problems before the customer is impacted — and use those learnings to drive product change, shape our internal tooling, and refine our operational processes so that we provide an increasing quality of service to more and more customers. Your approach is hands-on and pragmatic: you’ll rapidly address issues as they arise with quick and effective solutions, and advocate for workflow or product improvements once the immediate issue is resolved. You are energised by engaging directly with problems, from writing a script to automate a manual task, to finding creative workarounds, or building a case for a product enhancement. You don’t just fix issues — you look for opportunities to simplify, automate, and make the entire system more resilient. An FDRE synthesises learnings from support into best practices for others to follow. These are captured in documentation and shared with the team and the wider organisation. In this way, you raise the bar for reliability and efficiency across Palantir.

CV
Company via Lever
📍 New YorkFull-timeHybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Forward Deployed Reliability Engineer (FDRE), you ensure stability and reliability of mission-critical workflows built on Palantir software. You gather signal by going on call — resolving problems before the customer is impacted — and use those learnings to drive product change, shape our internal tooling, and refine our operational processes such that we provide an increasing quality of service to more and more customers. Your approach is hands-on and pragmatic: you’ll rapidly address issues as they arise with quick and effective solutions and advocate for workflow or product improvements after the immediate issue is resolved. You are energized by engaging directly with problems, from writing a script to automate a manual task, to finding creative workarounds, or building a case for a product enhancement. You don’t just fix issues— you look for opportunities to simplify, automate, and make the entire system more resilient. An FDRE synthesizes learnings from support into best practices for others to follow. These are captured into documentation and shared with the team and broader organization. In this way, you raise the bar for reliability and efficiency across Palantir.

🔔

Get new data reliability engineer jobs by email

Daily job updates · Unsubscribe anytime