Jobiba hiring network

Staff Software Reliability Engineer Data Platform Jobs

3,518 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current staff software reliability engineer data platform jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Work Flexibility: Onsite Stryker is seeking a Staff Advanced Manufacturing, Automation/Software Engineer to join our Advanced Operations team, supporting the Medical Division - Acute Care Business Unit. In this role, you will lead the design, development, and deployment of advanced automation and production test systems for new product introductions. This role is critical to ensuring reliable, scalable, and compliant manufacturing for patient support and patient environment products. You will serve as both a technical owner and a supplier-facing leader, architecting systems internally while guiding external partners to deliver high‑quality automation solutions. You will have the opportunity to work on the design transfer of new products from research through development and into production. This is a hybrid role based out of Portage, MI. The team works onsite 4-5 days per week to support collaboration and project needs. What you will do: Automation System Architecture & Development Lead the full lifecycle of industrial automation and test systems from requirements, architecture, and design through implementation, validation, and release. Define system-level requirements encompassing mechanical, electrical, software, controls, and safety considerations. Develop and integrate control software, embedded interfaces, test sequences, and operator interfaces (HMI/SCADA). Ensure robust performance, maintainability, reliability, and alignment with design intent and manufacturing needs. Troubleshoot complex processes, software, and equipment issues; optimize system performance and uptime. Supplier & Equipment Vendor Leadership Manage automation and equipment suppliers, including capability assessments, technical reviews, process monitoring, and on‑site visits. Create clear, comprehensive

pythonlinuxc#
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,

pythongcpdocker
View job →

Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu

pythonlinuxai
View job →
G
15 days ago

Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu

pythonlinuxai
View job →

Here’s a summary of the role: Build software that matters, take real technical ownership, and use modern AI tooling to do your best work. This is a hands-on senior engineering role for someone who enjoys solving complex product problems, shaping robust solutions, and helping teams deliver reliable services at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS, and modern engineering practices. You’ll play a leading role within a collaborative product engineering team, owning complex features end to end, contributing to design and architectural decisions, supporting production systems, and helping raise the bar across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Own and deliver complex backend services and APIs using Node.js, TypeScript, and AWS , from technical design through release and production support. Contribute to design and architecture discussions, making pragmatic decisions that balance delivery speed, maintainability, scalability, and security. Mentor and support less experienced engineers through code reviews, pairing, technical guidance, and day-to-day collaboration. Work closely with product managers, designers, and engineers across the team to turn requirements into practical, reliable solutions. Use AI tools to accelerate coding, debugging, testing, research, and documentation, while validating outputs carefully and applying sound judgment. Strengthen service reliability, observability, and engineering quality by improving monitoring, incident response, testing, and development practices. These are the essentials you’ll need to get an interview: 5 to 8 years of professional software engineering experience delivering production systems in an agile environment. Strong backend development s

typescriptreactnode.js
View job →

Here’s a summary of the role: Build cloud software that matters, grow your technical depth, and use modern AI tooling to do your best work. This is a hands-on engineering role for someone who enjoys solving product problems, writing clean code, and helping services run reliably at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS , and modern engineering practices. You’ll be part of a collaborative product engineering team where you can own features, contribute to design discussions, support production systems, and keep growing across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Design, build, test, and improve backend services and APIs using Node.js, TypeScript, and AWS . Take ownership of well-defined features from planning through release, including code quality, deployment, and production support . Work closely with product managers, designers, and other engineers to turn requirements into practical, reliable solutions. Contribute to technical design conversations, code reviews, and engineering standards that keep the team moving well. Use AI tools to speed up research, coding, debugging, testing, and documentation, while checking outputs carefully and applying sound judgment. Help keep systems secure, observable, and maintainable by improving monitoring, reliability, and day-to-day development practices. These are the essentials you’ll need to get an interview: 3 to 5 years of professional software engineering experience building production applications in an agile environment. Strong backend development skills with Node.js and TypeScript, including experience building APIs or microservices. Experience with React or Angular in a product engineering environment. Hands-on experience with

typescriptreactnode.js
View job →
F
Fin
📍 Ireland• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. We are an extremely product focussed team. We work in partnership with Product and Design functions of teams we support. Our team's dedicated ML product engineers enable us to move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure the real customer impact of each model we deploy. What will I be doing? Play an active role in hiring, mentoring and career development of other engineers Raise the bar for technical standards, performance, reliability, and operational excellence Identify areas

sqlrestmachine learning
View job →
F
Fin
📍 England• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. We are an extremely product focussed team. We work in partnership with Product and Design functions of teams we support. Our team's dedicated ML product engineers enable us to move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure the real customer impact of each model we deploy. What will I be doing? Play an active role in hiring, mentoring and career development of other engineers Raise the bar for technical standards, performance, reliability, and operational excellence Identify areas

sqlrestmachine learning
View job →
FB
15 days ago

Location : Come and join us in Hamburg. Freenow by Lyft empowers smarter mobility decisions helping people to move freely and cities to thrive. We are looking for an Engineering Manager to lead a high-performing team of software engineers focused on our Rider domain — the customer-facing app and platform that millions of riders use every day. Your team will play a pivotal role in bringing Lyft's product experience to Europe, adapting and building the systems that let riders book rides seamlessly across all markets This critical role ensures team’s alignment with product and business strategy, while driving efficient delivery of features. You will foster an environment of technical excellence, psychological safety, and continuous individual development, ultimately maximizing business impact across all Freenow markets through innovation. Be ready to work in a multinational, diverse, highly motivated and collaborative team of passionate colleagues who strive for excellence and enjoy their work. Are you ready for your next ride? YOUR DAILY ADVENTURES WILL INCLUDE: Team Performance & Delivery: Within the frame of quarterly goals and company strategy, lead the team in delivering high-quality, scalable software solutions. Focus on continuously improving engineering processes to ensure predictable, timely delivery while managing team capacity and focus. You inspire innovation to disrupt incumbents. People Development: Provide coaching, mentoring, and continuous feedback to team members. Identify development areas and growth opportunities to support individual career paths and enhance team competence in building innovative customer and driver support solutions. You get to build a team from the ground up. Technical Ownership: In collaboration with Staff Engineers and Principals, lead technical decision-making for the support systems. Ensure high standards of quality, reliability, maintainability, and observability in all solutions, particularly the large-scale syste

javaawsdocker
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Core Infrastructure team's mission is to build and evolve the foundational platform that every Robinhood engineering team builds on — owning the systems, primitives, and developer-facing abstractions that power 24/7 trading, crypto, and global expansion. We treat infrastructure as a product: reliable, fast to provision, and invisible to the teams above it. As a Senior Staff Software Developer on Core Infrastructure, you will own the architectural evolution of three deeply interconnected domains: service mesh and connectivity, compute platform, and infrastructure provisioning. Your decisions will directly shape engineering velocity, operational reliability, and Robinhood's ability to expand to new regions and markets. This is not an operations role — it's a once-in-a-platform-lifecycle opportunity to redesign the foundation before complexity becomes permanent! This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-p

pythonkubernetesartificial intelligence
View job →
G
Gitlab
📍 United States• Full-time• Remote• From $126.4K/yr
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Staff Human Resources Information Systems Analyst, People Technology - Workday, you won't just manage our people systems—you'll own them. You'll drive how our people technology evolves, partner deeply with stakeholders to solve root problems, and build scalable solutions that reduce friction and improve reliability, usability, and insight across the People technology landscape. This is a strong fit if you think like a product owner, refuse to accept requests at face value, and can move with speed and precision within public company compliance requirements, including SOX ITGCs. In this role, you'll le

REMOTEgitrestai
View job →
E
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i

javaawskubernetes
View job →
E
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Portworx team to build and deliver our highest-quality product suite. In this role, you will write clean, scalable code with a strong focus on quality, reliability, and user-centric design. You will directly contribute to building a new SaaS platform that delivers a secure, consistent, and best-in-class experience for customers purchasing and managing Portworx offerings. As a core developer, you will take ownership of designing and implementing critical features across the entire Portworx portfolio. WHAT YOU’LL DO Design & Scale SaaS Microservices: Develop, test, and integrate high-performance microservices and features into the Portworx product suite, ensuring high availability in distributed systems. Drive End-to-End Delivery: Lead software lifecycle activities including architectural design, code reviews, unit/functional testing, documentation, and continuous integration and deployment (CI/CD). Partner Across Teams: Collaborate with product managers, cross-functional engineering peers, and early-adopter customers to transform requirements into production-ready software. Own Product Quality & Iteration: Take full ownership of feature stability by proactively incorporating customer feedback and rapidly resolving issues identified during testing and deployment. Innovate & Experiment: Research emerging technologies and cloud infrastructure tools to push performance boundaries and continuously i

javaawskubernetes
View job →
E
Everpure
📍 Bengaluru• Full-time
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE The FlashArray team builds an innovative, high-performance, and highly available portfolio of products designed for demanding, mission-critical applications. While we deliver a hardware storage array, over 90% of our engineering team focuses on software engineering. Our customers value FlashArray for its simplicity of management, continuous feature upgrades, and ability to stay on the cutting edge without downtime. Recently, we extended these capabilities into the public cloud with CloudSnap and Cloud Block Store for AWS, enabling customers to leverage cloud agility for both traditional IT and cloud-native applications. WHAT YOU’LL DO Design & Implement: Create innovative algorithms and technologies for high-performance systems targeting six-nines (99.9999%) reliability. End-to-End Ownership: Lead feature innovation from initial concept through to shipped product. Problem Solving: Analyze and resolve complex technical challenges through persistent problem-solving and technical insight. Collaborate & Deliver: Partner closely with smart, collaborative peers to deliver features that directly enhance customer experience and satisfaction. Learn & Grow: Continuously expand domain expertise in systems software within a supportive, knowledge-sharing environment. WHAT YOU BRING Software Development Experience: 3+ years of professional development experience using C, C++, Python, Go, Java, or related prog

pythonjavaaws
View job →
Z
Zscaler
📍 Bellevue• Full-time• From $180K/yr
15 days ago

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Senior Staff Rust Developer to join our Platform Convergence Team. This is a hybrid role based in San Jose, CA reporting to the Sr. Director, Software Engineering. Join us to build a new platform from the ground up that can scale hundreds of millions of users with high reliability and low latency. You will design and implement distributed system and core infrastructure components while collaborating closely with various stakeholders. What you’ll do (Role Expectations) Design and build a low-latency, high-throughput data forwarding plane using Rust, leveraging its async/await model for efficient I/O and service-oriented infrastructure Develop distributed, scalable systems with a focus on concurrency, fault tolerance, and messaging Implement and maintain gRPC-based APIs and services to integrate forwarding plane capabilities with control and orchestration layers Optimize system

awskubernetesci/cd
View job →
🔔

Get new staff software reliability engineer data platform jobs by email

Daily job updates · Unsubscribe anytime