Jobiba hiring network

Software Reliability Engineer Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

About the Team The SaaS and Software Governance team sits within Corporate IT and helps OpenAI scale from startup-speed tooling to mature enterprise architecture. The team owns practical governance for software, SaaS, integrations, APIs, identity, data access, and agent-enabled workflows, with a mandate to improve security, reduce software sprawl, and help business teams move faster through better foundations. About the Team OpenAI is scaling from an emerging, high-velocity startup into a mature enterprise operating model. The IT Software Architect will help shape the software, SaaS, Data integration, and agent-enabled architecture that lets the company move quickly while improving security, compliance, data quality, and customer, partner, and employee experience. This role is not a traditional ivory-tower architecture function. It is a hands-on governance and enablement role that partners with business teams, IT operations, Security, Procurement, Business Platforms, Applied teams and Data Engineering to guide software decisions, reduce unmanaged sprawl, and build reusable enterprise foundations. Why This Role Matters OpenAI’s software footprint is expanding rapidly across SaaS, internally built tools, agents, integrations, APIs, third-party platforms, and application systems that OpenAI. The company needs a stronger tools architecture layer that can help teams make good decisions early, avoid duplicate tools, govern sensitive data and identities, and identify where OpenAI should build instead of buy. The person in this role will help turn software governance from an approval checkpoint into an enterprise capability: a system that improves speed, reliability, security, and business outcomes. What You'll Do Own the target architecture for enterprise software, SaaS, integrations, APIs, and agent-enabled business systems across Corporate IT. Drive deprecation and consolidate targets for enterprise software. Build lightweight governance patterns that guide teams before

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior Android engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable Android foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Have a proven track record of building high-quality Android applications in production. Are fluent in Kotlin (and/or Java) and familiar with Android development tools and architecture components. Prioritize performance, security, and user experience in mobile development. Enjoy working cross-functionally to bring ambitious product ideas to life. Care deeply about performance, security, and user experience. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We

javaawsrest
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior iOS engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable iOS foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. In this role, you will: Build and ship new experiences on iOS that showcase the power of AI. Optimize app performance, reliability, and responsiveness at global scale. Design and maintain shared iOS frameworks and primitives for account, trust, and commerce flows that are used across OpenAI’s mobile apps. Establish robust testing frameworks and refine app architecture for long-term maintainability. Collaborate with product, design, research, and backend teams to deliver high-impact features. Provide technical leadership to shape the future of OpenAI’s iOS platform. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Hav

awsrestai
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring a talented Software Engineering Manager to lead the Snowtrail infrastructure team at Snowflake. Snowtrail is the infrastructure that enables Snowflake to deliver dedicated coverage for customer-specific workloads. Its innovative approach allows Snowflake to precisely test and measure the impact of changes on individual customers, making it essential for ensuring the platform’s reliability, correctness, and performance. Through query replay, Snowtrail helps us catch regressions early. By leveraging machine learning models to intelligently sample queries and workloads, we continuously optimize for both cost and performance. Evolving Snowtrail to incorporate new engine features while improving scalability, efficiency, and reliability is central to our continued success OUR IDEAL MANAGER WILL HAVE : Strong passion and proven track record for shipping quality software in high code velocity environments 10+ years industry experience designing and building distributed data systems. Excellent problem solving skills, and strong CS fundamentals including data structures, algorithms, and distributed systems. Fluency in SQL, Java, C++, Python or Go. Ability to collaborate well across teams, build high-performing teams and mentor junior engineers. Excellent interpersonal co

pythonjavavue
View job →
E
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Drive the quality strategy for our innovative enterprise storage platform, ensuring zero-downtime resilience across physical hardware and cloud environments like Cloud Block Store and CloudSnap. In this engineering leadership role, you will scale systems testing, feature interoperability, and test automation for mission-critical global applications. Partnering directly with cross-functional development, support, and escalation teams, you will champion a customer-first quality model. This position elevates overall product reliability while shaping how cutting-edge software resilience is delivered at scale. WHAT YOU'LL DO Define & Execute Quality Strategy: Own end-to-end system test designs with a focus on large-scale feature interoperability to guarantee zero-downtime performance across enterprise and cloud environments. Build High-Impact Automation & Tooling: Design and deploy automated test workflows and triage tooling to accelerate defect detection, drastically reducing execution friction across thousands of automated test suites. Simulate Real-World Customer Workflows: Replicate complex customer deployment architectures to validate real-world fault tolerance and overall resilience against failure domains. Drive Root-Cause Resolution: Partner directly with escalation and support engineering teams to analyze and resolve complex defects, utilizing customer feedback loops to eliminate quality gaps. Lead Agi

javaawsrest
View job →
E
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Drive the mission-critical quality strategy for the industry’s most innovative, high-performance storage array platform. In this pivotal engineering leadership role, you will scale systems testing, feature interoperability, and test automation to ensure zero-downtime resilience for global enterprise applications. Partnering directly with cross-functional development, support, and escalation engineering teams, you will champion a customer-first quality model across physical hardware and cloud-native environments (Cloud Block Store, CloudSnap). This position elevates product reliability and shapes how cutting-edge software resilience is delivered at scale. WHAT YOU'LL DO Define & Execute Quality Strategy: Ownership of end-to-end system test designs, focusing on feature interoperability at scale to guarantee zero-downtime performance across enterprise and cloud environments. Build High-Impact Automation & Tooling: Design and deploy automated test workflows and triage tooling to accelerate defect detection, drastically reducing execution friction across thousands of automated test suites. Real-World Customer Simulation: Replicate complex customer deployment architectures and enterprise application workflows to validate real-world resilience, fault tolerance, and resilience against failure domains. Root-Cause Resolution & Continuous Improvement: Partner directly with escalation and support teams to reproduc

javaawsrest
View job →
E
Everpure
📍 Prague• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As the Engineering Manager for FlashArray File, you will lead a Prague-based team dedicated to evolving Everpure’s™ industry-leading file services. You’ll drive the mission of delivering enterprise-grade, high-availability file capabilities that empower global innovators to manage mission-critical data. This role differentiates itself by blending deep systems-level engineering with high-impact people leadership, requiring close collaboration with Product Management and global R&D teams to redefine the modern data experience. Your goal is to ensure our customers can seamlessly scale their file environments with the reliability and performance Everpure™ is known for. WHAT YOU’LL DO Drive End-to-End Execution: Lead the full software development lifecycle (SDLC) for core file features, ensuring the team delivers high-quality, scalable code from initial design through production rollout. Shape Technical Strategy: Partner with senior architects and Product Management to define the roadmap for multi-server and directory services, making critical trade-offs that balance innovative feature delivery with long-term system stability. Cultivate Engineering Excellence: Champion rigorous testing strategies and CI/CD health, fostering a "follow-your-feature" culture where engineers take ownership of operational excellence and defect prevention. Empower and Develop Talent: Manage and mentor a team of 5–8 engineers, providing t

awsci/cdrest
View job →

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Core Infrastructure team's mission is to build and evolve the foundational platform that every Robinhood engineering team builds on — owning the systems, primitives, and developer-facing abstractions that power 24/7 trading, crypto, and global expansion. We treat infrastructure as a product: reliable, fast to provision, and invisible to the teams above it. As a Senior Staff Software Developer on Core Infrastructure, you will own the architectural evolution of three deeply interconnected domains: service mesh and connectivity, compute platform, and infrastructure provisioning. Your decisions will directly shape engineering velocity, operational reliability, and Robinhood's ability to expand to new regions and markets. This is not an operations role — it's a once-in-a-platform-lifecycle opportunity to redesign the foundation before complexity becomes permanent! This role is based in our Toronto, ON office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-p

pythonkubernetesartificial intelligence
View job →

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Rider Loyalty team is where riders become members. We build the membership, rewards, and benefits products that give people a reason to choose Lyft on every trip, and we make sure the value a rider has earned shows up at the moment it matters. Loyalty sits inside the Rider Loyalty, Partnerships, and Rider Pay (PLP) group. You will lead a team of engineers across iOS, Android, and Server. You will own the membership and rewards platform end to end and work daily with Product, Design, Data Science, and Partnerships. Responsibilities: Own the Loyalty roadmap from strategy through delivery. Turn goals like member growth and retention into an engineering plan, and manage the dependencies that run through Partnerships and Rider Pay. Build and scale the systems behind membership, rewards earning and redemption, and benefit delivery. Hold a high technical bar through architecture reviews, tech debt management, observability, reliability, and on-call. Grow engineers by matching people to the right opportunities, setting clear expectations, and giving feedback early. Experience: 5+ years building software professionally, including 2+ years directly managing engineers. You have managed a team that shipped both mobile and backend work, and you can still read and review code in at least one of those areas. You have owned a consumer product used by millions of people each month. You use AI tools in your own work and have a clear view of where they help and where they do not. BS/MS in Computer Science, Computer Engineering, or a related field, or equivalent practical experience. Benefits: Extended health and dental coverage options, along with life insurance and disability benefits Mental health benefits Family building benefits Child care and pet benefits Access to a Lyft funded Health Care Savings Accou

artificial intelligenceai
View job →
C
Clickup
📍 Canada• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ

node.jsawsmachine learning
View job →
C
Clickup
📍 United States• Full-time
1mo ago

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 ClickUp is looking for an experienced Engineering Manager to lead our fullstack team responsible for building and scaling our flagship products. As the leader of the team that owns the APIs and core experiences powering ClickUp, you will play a pivotal role in shaping the future of our platform. You will guide engineers working across the stack, from frontend experiences to backend infrastructure. Your focus will be on driving the development of new features, addressing performance and reliability challenges, and ensuring operational excellence as we continue to grow. This is an opportunity to make a significant impact on our core product while fostering a culture of technical excellence and collaboration. The Role: Technical Leadership : Provide hands-on technical guidance to the team, ensuring best practices in software development, architecture, and design. Team Management : Lead, mentor, and grow a team of engineers, fostering a culture of collaboration, innovation, and continuous improvement. Product Development : Drive the development of new features and enhancements, ensuring high performance, scalability, and reliability. Collaboration : Work closely with product managers, designers, and other engineering teams to align on goals, prioritize initiatives, and deliver exceptional user experiences. Code Quality : Oversee code reviews, ensure adherence to coding standards, and advocate for clean, maintainable, and testable code. Innovation : Stay up-to-date with the latest trends and technologies in collaborative editing, cloud infrastructure, and web development, and apply them to improve our produ

node.jsawsmachine learning
View job →

Observability Pipelines (OP) is Datadog's on-premise, vendor-agnostic telemetry pipeline product. As an Engineering Manager on the team, you'll own people management and engineering execution for one of OP's core missions, spanning areas like Integrations (ingesting from and routing to the many source and destination systems customers rely on), streaming insights, cost control, or pipeline capabilities, reliability and scalability. You'll partner directly with Product to help shape the roadmap, and work closely with your peer EMs and senior ICs to define how OP operates and grows. This is an opportunity to build your management craft while having real influence over the technical direction of a fast-growing product area. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Own people management and engineering execution Establish a strong operating rhythm for the team Drive high standards for on-call rotations and incident response Partner with Product on the roadmap, balancing product priorities with technical realities Lead, coach, and grow the careers of engineers on your team Who You Are: Experienced managing engineers directly, comfortable owning a team’s operating rhythm end-to-end, from planning through execution and stakeholder communication to incident and on-call ownership Have a technical background in distributed systems and data infrastructure Have experience with on-premises or customer-installed software concepts A product-minded partner to have on the team — you enjoy working with Product on strategy Experience with high-performance or Rust-based data pipeline systems Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications o

aigorust
View job →
O
1mo ago

About the Team The Finance & Supply Chain Engineering organization includes two complementary teams. Software Engineering builds internal full-stack applications, durable agentic workflows, plugins, MCPs, and measurable AI-enabled engineering practices. Data Engineering builds trusted analytics data assets for Finance and Supply Chain. The teams have distinct charters, with important shared dependencies and broad cross team partnerships across Engineering, Applications, Finance, and Supply Chain. About the Role We are looking for a hands-on senior technical leader who will report alongside the Software Engineering and Data Engineering managers. This is an individual-contributor role with no immediate people-management responsibility. The Tech Lead will raise the technical bar across both teams, participate in important cross-team or high-risk design decisions, and directly own and ship high-impact work. The role should improve team judgment and autonomy rather than act as a floating architect or universal approval gate. In this role, you will: Partner with the Software Engineering and Data Engineering managers as a peer technical leader; managers retain accountability for people, staffing, priorities, performance, and delivery commitments. Directly own the architecture, implementation, launch, and operation of one or more high-impact initiatives, remaining accountable for real outcomes rather than advisory output alone. Guide important design decisions that are cross-team, difficult to reverse, or material to security, financial controls, reliability, data quality, or long-term cost of ownership. Establish pragmatic engineering standards across architecture, APIs and data contracts, testing, security, reliability, observability, lineage, data quality, and operational ownership. Advance engineering standards for building with AI, including agentic workflows, evaluation, telemetry, adoption, and outcome measurement. Work across backend, full-stack, data, and big-d

awsrestai
View job →
E
Everpure
📍 Bengaluru• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As an Escalation Engineer on our FlashBlade team, you will serve as the premier technical authority driving customer trust and operational stability across complex, enterprise-scale storage environments . You will collaborate closely with front-line Support, Engineering, and product leaders to rapidly resolve high-impact technical challenges and transform complex system failures into long-term product reliability. By bridging real-world customer insights with engineering solutions, you will elevate team performance and ensure our enterprise customers achieve flawless platform availability. WHAT YOU'LL DO Drive High-Stakes Escalation Resolution: Take end-to-end ownership of critical, multi-platform system issues—evaluating hardware, software, networking, and environmental factors—to rapidly restore service, perform root-cause analysis, and protect customer business continuity. Elevate Engineering Talent & Knowledge: Mentor and coach support team members through joint case triage, structured technical trainings, and internal documentation, accelerating technical capabilities and resolution velocity across the organization. Bridge Product Engineering & Customer Insights: Partner directly with Product Engineering to relay real-world system behavior, ensuring critical customer feedback, feature enhancements, and bug fixes trickle back into core product design. Lead Strategic Customer Communications: Facilitate

awslinuxrest
View job →
E
Everpure
📍 Bengaluru• Full-time
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As an Escalation Engineer on our FlashBlade team, you will serve as the premier technical authority driving customer trust and operational stability across complex, enterprise-scale storage environments . You will collaborate closely with front-line Support, Engineering, and product leaders to rapidly resolve high-impact technical challenges and transform complex system failures into long-term product reliability. By bridging real-world customer insights with engineering solutions, you will elevate team performance and ensure our enterprise customers achieve flawless platform availability. WHAT YOU'LL DO Drive High-Stakes Escalation Resolution: Take end-to-end ownership of critical, multi-platform system issues—evaluating hardware, software, networking, and environmental factors—to rapidly restore service, perform root-cause analysis, and protect customer business continuity. Elevate Engineering Talent & Knowledge: Mentor and coach support team members through joint case triage, structured technical trainings, and internal documentation, accelerating technical capabilities and resolution velocity across the organization. Bridge Product Engineering & Customer Insights: Partner directly with Product Engineering to relay real-world system behavior, ensuring critical customer feedback, feature enhancements, and bug fixes trickle back into core product design. Lead Strategic Customer Communications: Facilitate

awslinuxrest
View job →
🔔

Get new software reliability engineer jobs by email

Daily job updates · Unsubscribe anytime