Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. About the Role & Team Everpure acquired Portworx to create the industry's most complete Kubernetes Data Services Platform for cloud-native applications. In this role, you will be supporting multi-cloud data services for Kubernetes and containerized workloads, enabling leading enterprises to run mission-critical data applications smoothly across public and private clouds. WHAT YOU'LL DO Analyze & Support: Troubleshoot and support large-scale customer deployments in public/private clouds across all severity levels. Technical Expertise: Provide hands-on guidance during all deployment phases, POCs, pre-sales calls, and production environments for key accounts. Cross-Functional Collaboration: Partner with engineering teams to analyze logs, reproduce complex customer issues, and develop long-term fixes. End-to-End Ownership: Track customer support cases end-to-end, triage multi-layer software stack issues, and escalate to core engineering when needed. Knowledge Sharing: Author and maintain KB articles, FAQs, and technical documentation for internal teams and customers. WHAT YOU BRING Experience: 2 - 4+ years in customer-facing technical support or Site Reliability Engineering (SRE). Containers & Orchestration: Solid working knowledge of Kubernetes, OpenShift, Tanzu, or VMware container solutions ( CKA certification is a plus ). Cloud Platforms: Hands-on experience with AWS, Azure, GCP, or related cloud technol
Jobiba hiring network
Reliability Engineer Jobs
2,049 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Automation is the key to creating highly reliable and secure large-scale software systems. Are you someone who engineers solutions to problems rather than simply fixing the same thing over and over again? Can you protect Smartsheet against attackers? We are looking for a Senior DevSecOps Engineer to join our global Security Operations team. In this critical role, you will be a leader in maturing our security and reliability posture by treating both as software engineering challenges. You will engineer and operate a highly reliable, scalable, and defensible production environment, directly impacting our ability to deliver a world-class service to our customers 24/7. This is a unique opportunity to blend deep expertise in Site Reliability Engineering (SRE) and modern Security Operations, working at the intersection of infrastructure, automation, and security to build a platform that is resilient and secure by design. You Will: Engineer Secure and Resilient Infrastructure: Design, build, maintain, and improve secure, scalable, and highly available infrastructure in our multi-cloud environment (primarily AWS) using Infrastructure as Code (IaC) principles with tools like Terraform, Kubernetes, and Helm. Automate Proactive Security: Engineer an
About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage international teams. With our core values at heart and future focused work culture, our team works tirelessly on ambitious problems, asynchronously, around the world. You can find Remoters working from 6 different continents (Antarctica left to go!) and all of our positions are fully remote. With Innovation as one of the core values, we have built Automation and AI capabilities into the requirements for every role. We encourage every member of the Remote team to bring their talents, experiences and culture to the table to help us build the best-in-class HR platform. If you are energetic, curious, motivated and ambitious, be part of our world. Apply now and define the future of work! What this job can offer you Remote's SRE team exists so that our engineers can move quickly and our customers get a product that stays up. The team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, our observability stack and the reliability practices that sits on top of all of it. We are looking for a Team Leader to run that team. This is a 60% IC, 40% leadership role. You will own the career development of your reports, steer the teams focus using judgment against the company goals, and you will be the spokesperson for the team across engineering. You will also stay close enough to the technical work to set direction with credibility and to know when something is going wrong before it is escalated to you. Reliability practice at Remote is maturing rather than mature. Our SLO framework is live on its first few teams and needs to reach the rest, there is real work to do on how we balance operational load against project delivery. If you want a team where the foundations are in place and the interesting problems are still open, this is that team. What you bring People leadership Yo
DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities. Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos. About the Role We're seeking an experienced DevOps/ Site Reliability Engineering (SRE) Engineer to join DataHub and drive the reliability, scalability, and operational excellence of our platform offerings. In this role, you'll work on technical initiatives across DataHub Cloud and our emerging enterprise deployment solution, which provides customers with enhanced control and flexibility for running DataHub in their preferred environments. Key Responsibilities Enterprise Platform Development: Partner with product and engineering teams to influence the development of advanced deployment capabilities. Collaborate with cross-functional teams to help build systems for seamless installation, upgrade, and rollback processes across various environments. Influence the design and help implement comprehensive monitoring and health check systems for distributed deployments. Partner with engineering teams to help develop self-healing and automated remediation capabilities. Platform Reliability and Operations: Establish and maintain SLAs/SLOs for both cloud and enterprise offerings. Lead incident response and post-mortem processes to drive continuous improvement. Optimise system performance, capacity planning, and cost efficiency. Work closely with product, engineerin
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, http://incident.io , Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different
TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. T extNow is looking for a new SecOps team member to secure, monitor , and enable automated response within our infrastructure. What You’ll Do Ensure Secure & Reliable Systems: Design, implement, and maintain security-focused infrastructure to protect TextNow’s services while ensuring reliability and scalability. Security Automation & Infrastructure as Code: Develop and enforce best practices using Terraform, Ansible, Crowdstrike , and AWS security tools , ensuring secure configurations, automated compliance checks, and infrastructure as code. Threat Detection & Incident Response: Participate in an on-call rotation to respond to security incidents, investigate vulnerabilities, and implement proactive measures to prevent future threats. Work closely with engineering teams to remediate security risks. Monitoring & Logging for Security: Improve observability by implementing security monitoring solutions, logging best practices, and alerting mechanisms to detect anomalies and suspicious activity. Access Control & Identity Management: Manage IAM roles, permissions, and policies to ensure least privilege access and enforce security controls across cloud and internal systems. Collaboration & Security Advocacy: Wo
We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s
About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa
Job Details: Job Description: As a Material Analysis (MA) Technician, you will be part of a Technology Development (TD) and High-Volume Manufacturing (HVM) lab responsible for performing material analysis and failure analysis in support of Intel's silicon process development and high-volume production. You will work on developing imaging, composition analysis and sample preparation techniques, and best-known methods (BKMs) to improve lab analysis quality, efficiency and output. You will directly interface with TD and HVM fab customers and quality/reliability engineers to develop solutions to problems by utilizing lab capabilities. The scope may include wafer and unit level, front-end modules and back-end/far back-end modules. Responsibilities may include but not be limited to: • Conducting hands-on analysis by effectively utilizing lab techniques, from sample prep micro-cleaver, ion mill etcher, mechanical polish to SEM/EDX, Dual beam FIB, TEM techniques to characterize Si fabricated structures at nanometer scales and integrated circuit device to improve process, performance and reliability; and to identify physical failure mode toward the root-cause identification. • Conducting hands-on data collection with various lab equipment’s, and assisting engineers to implement materials characterization techniques to determine fundamental thin film material structure/properties and to collaborate with process development engineers across functional areas and organizations to improve process performance and reliability. • Supporting and sustaining lab equipment. Ensuring that lab analytical capabilities needed to support advanced transistor and interconnect technology and/or product development are in place. Cooperating with other lab areas beyond local MA/FA (Failure Analysis) labs to achieve
About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late
Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. mthree has an exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input
Want to work in technology at an investment bank? Paid graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. mthree has an exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user
Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches).Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel.This is an exciting opportunity to join an FX Front Office support team in a major North American bank working on the Toronto Trading floor, supporting both front office users and a progressive eFX programme. What you'll do: Support IT solutions for various business lines globally including FX, Money Market, STIR and Options Offer technical expertise and support for the systems used Manage and resolving incidents and outages Manage requests for changes, system releases and capacity planning Disaster recovery, planning and execution How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 9 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input, common data structures, and flow of control. How to use MySQL to perform CRUD (create, read, update and delete) operations on a relational database stored in
Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches).Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel.This is an exciting opportunity to join an FX Front Office support team in a major North American bank working on the Toronto Trading floor, supporting both front office users and a progressive eFX programme. What you'll do: Support IT solutions for various business lines globally including FX, Money Market, STIR and Options Offer technical expertise and support for the systems used Manage and resolving incidents and outages Manage requests for changes, system releases and capacity planning Disaster recovery, planning and execution How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 9 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input, common data structures, and flow of control. How to use MySQL to perform CRUD (create, read, update and delete) operations on a relational database stored in
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime