At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Our team builds and operates Snowflake's Streaming Platform — the core infrastructure responsible for bringing data into Snowflake continuously and at scale. This includes Snowpipe Streaming and Datastream, services used by large set of enterprise customers to move mission-critical data in real time. Data is the fuel powering the new enterprise AI world, and we are the team responsible for getting it there reliably, continuously, and fast. Responsibilities Design, build, and maintain core components of Snowflake's streaming ingestion platform Improve service reliability, scalability, and latency under high-throughput production workloads Debug and resolve incidents in a distributed, multi-tenant cloud service Contribute to the design and evolution of streaming APIs, SDKs, and server-side protocols Write comprehensive tests including unit, integration, and chaos/fault-injection scenarios Collaborate with partner teams (storage, query, compute) on cross-cutting platform concerns Participate in code reviews, on-call rotations, and architecture discussions Required Qualifications 3–5 years of software engineering experience on large-scale distributed systems or cloud services Strong proficiency in Java or C++ Deep understanding of distributed systems concepts: consistency, faul
Jobiba hiring network
Reliability Engineer Jobs
2,028 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. AS A SENIOR SOFTWARE ENGINEER YOU WILL: Drive high-impact initiatives that span our product areas and full tech stack, including golang and Python on the backend and TypeScript/React on the frontend. Own and deliver features across the notebook service, container runtimes, and UI — designing and shipping medium-to-large projects independently, from ambiguous problem statements through production and post-launch. Advance core platform initiatives such as runtime management and patching, environment reproducibility and replication, security and compliance, and observability for notebooks. Extend the product to operate reliably in regulated and air-gapped environments, where security, compliance, and operational rigor are paramount. Promote strong collaboration within a cross-functional team and partner closely with embedded product managers and designers, as well as platform organizations across Snowflake Be a strong contributor to the product vision and drive team planning. Build for scale, reliability, and high performance, and participate in the on-call rotation to keep a Tier-1 production service healthy. Mentor, coach, and empower more junior team members, and raise the engineering bar through high-quality design and code review. OUR IDEAL CANDIDATE WILL HAVE: 7+ years o
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. We are hiring a Senior Frontend Engineer, AI Products team at Observe by Snowflake. As an AI Product engineer you'll always be thinking first about the user experience and how to create the best product, technical choices, and implementation decisions that stem from that product first thinking. This team builds the AI-powered products and developer tooling at the core of Observe's platform, including our flagship AI SRE product, real-tim
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. As a Senior Technical Support Engineer, you will be a trusted advisor and technical resource for our customers in the EMEA region. This is a hands-on role for someone who thrives in dynamic environments, loves troubleshooting complex technical issues, and is passionate about delivering exceptional support experiences. You’ll be responsible for resolving high-impact technical issues, driving customer success, and collaborating closely wit
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for the Observe Data Management team. This team owns the core pipelines that ingest and process over 1 petabyte of telemetry data per day — the foundational infrastructure powering Observe's entire observability stack. You'll be working at the intersection of massive scale, open-source innovation, and real-world reliability challenges for enterprise customers around the globe. AS A SENIOR SOFTWARE ENGINEER - OBSERVE
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Customer Experience Engineering Team builds the internal and external technologies that scale Snowflake’s global support and sales organizations. We empower our technical experts by providing the advanced tools they need to resolve complex issues and drive customer success. Our team specializes in software engineering, data-driven decisions, ML, and LLM-based solutions . We build production-grade systems to automate manual processes and augment the capabilities of our technical staff. Our current focus includes: LLMs : Developing and deploying LLM and agent-based architectures for streamlining troubleshooting Scalable Evaluations : Implementing large-scale evaluations to ensure the quality and reliability of our internal and external tools Process Automation : Designing intelligent workflows that eliminate bottlenecks and allow our experts to focus on the most technical aspects of the Snowflake platform Incident discovery: using embeddings, LLMs, clustering, and agents to detect potential widespread issues more quickly Now, the team is growing, and we are looking for a Software Engineer to join us. In this role, you will work closely with the state of the art LLM models, fine-tune them, develop agents, apply various clusterings, summarizations, embeddings, and so on. Ev
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the team The Apps & Experiences Platform team is powering systems and services for all our user facing apps including Snowsight , Snowflake Intelligence , and new mobile apps . Our mission is to craft innovative backend services, features, tools, infrastructure, and AI tooling that bring such products to life with delightful user experiences. As part of our team, you'll dive into a mix of building & managing platform infrastructure, and building AI self-serve tools to support the platform and its developer’s needs. We're passionate about building a platform that is highly reliable, available, maintainable, and scalable. We are a high growth AI data cloud company and we are looking for exceptional talent like you to help build and grow our infrastructure to scale us to the next level. AS MANAGER FOR APPS & EXPERIENCES PLATFORM TEAM, YOU WILL: Own the technical strategy and execution for the team, driving projects from initial idea formulation and detailed system design to high-quality implementation and successful deployment. Provide deep technical oversight by actively participating in design reviews, architecture discussions, and drilling into complex system implementations to ensure reliability and scalability. Serve as a subject matter expert , setting
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring a talented Software Engineering Manager to lead the Snowtrail infrastructure team at Snowflake. Snowtrail is the infrastructure that enables Snowflake to deliver dedicated coverage for customer-specific workloads. Its innovative approach allows Snowflake to precisely test and measure the impact of changes on individual customers, making it essential for ensuring the platform’s reliability, correctness, and performance. Through query replay, Snowtrail helps us catch regressions early. By leveraging machine learning models to intelligently sample queries and workloads, we continuously optimize for both cost and performance. Evolving Snowtrail to incorporate new engine features while improving scalability, efficiency, and reliability is central to our continued success OUR IDEAL MANAGER WILL HAVE : Strong passion and proven track record for shipping quality software in high code velocity environments 10+ years industry experience designing and building distributed data systems. Excellent problem solving skills, and strong CS fundamentals including data structures, algorithms, and distributed systems. Fluency in SQL, Java, C++, Python or Go. Ability to collaborate well across teams, build high-performing teams and mentor junior engineers. Excellent interpersonal co
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. In this role you will: Develop interactive, data-rich user interfaces using React, TypeScript, and Vega, with a focus on integrating LLM-driven features (e.g., natural language querying, generative UI, and AI-assisted data storytelling). Lead the end-to-end delivery of substantial product features, ensuring AI outputs are presented with high reliability and low latency. Work closely with PMs, UX designers, and AI/ML engineers to bridge t
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. The Role You will work on our Metrics platform - enabling our users to query, visualize, and alert on billions of time series quickly and effectively. You'll own meaningful parts of that stack, drive performance and scalability improvements, and contribute to architectural decisions that shape where the platform goes next. This isn't a maintenance role, the metrics backend is being actively evolved, and you'll be a core part of that. Thi
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse, using open formats like Apache Iceberg, at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause of production issue and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world’s leading data platforms. The Role You'll be our dedicated expert in query execution and query performance. That means owning the query execution service end-to-end: working on caching strategies, incremental execution, query rewrites, and other optimizations that directly affect the speed and cost of running Observe at scale. You'll also be the go-to resource when query latency issues arise during customer evaluations and new deal cycles, diagnosing root causes
Job Title: Senior QA Engineer - Performance Testing Paytm is India's leading mobile payments and financial services distribution company. A pioneer of the mobile QR payments revolution in India, Paytm builds technologies that empower small businesses with payments and commerce solutions. Paytm’s mission is to serve half a billion Indians and bring them into the mainstream economy through the power of technology. About the Role: We are seeking a skilled Performance Test Engineer to design, execute, and analyze performance tests to ensure application scalability, stability, and responsiveness under varying load conditions. The ideal candidate will have hands-on experience with industry-standard performance testing tools and a strong understanding of system architecture, monitoring, and troubleshooting. Expectations/ Requirements Develop comprehensive performance test strategies and plans aligned with system requirements, project timelines, and business goals. Understand application architecture and identify critical business transactions for performance validation. Design realistic workload models to simulate real-world usage scenarios. Create, maintain, and execute performance test scripts using tools such as JMeter, LoadRunner, Gatling, or similar. Conduct baseline, load, stress, and scalability testing to evaluate system behavior under different conditions. Monitor system performance using tools like Influx DB, Grafana, JVM monitoring tools, and MAT (Memory Analyzer Tool). Analyze test results to identify performance bottlenecks and system limitations. Collaborate with development and infrastructure teams to troubleshoot and resolve performance issues. Assess system scalability and recommend optimizations to improve performance and reliability. Generate detailed performance test reports, including metrics, findings, and actionable recommendations. Work with stakeholders to gather and validate Non-Functional Requirements (NFRs), SLAs, and KPIs. Perform API an
About Us: Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. About Role: We are seeking an experienced L2 Network & Security Engineer to join our team. As an integral part of our network operations, you will play a crucial role in maintaining and securing our infrastructure. If you have a passion for networking, security, and troubleshooting, we’d love to hear from you! Job Location: Noida Responsibilities: Install and Support: Deploy and maintain LANs, WANs, and network segments. Support Coordination: Collaborate with equipment vendors for troubleshooting and configuration standardization. Firewall Management: Handle firewall configurations to ensure security and optimal performance. Switch Expertise: Maintain a strong understanding of LAN switching technologies, including VLANs, RSTP, ACLs, and Virtual Chassis. Capacity Planning: Recommend network capacity planning strategies. Proactive Maintenance: Plan and execute proactive maintenance for disaster recovery. Network Monitoring: Monitor networks for security, reliability, and availability. Key Skills Required: He/She/They should have 3-6 Yrs. of overall IT experience. Firewalls: Proficient in configuring, managing, and troubleshooting FortiGate and Palo-Alto firewall. Network Switches: Skilled in handling Cisco, Juniper, and Huawei switches. Wireless: Familiarity with Aruba, Cisco controllers and access points. Cloud Expertise: Experience with AWS and Zscaler Private Access. Network Protocols: Sound knowledge of various network protocols and ports. Security Technologies: Familiarity with IPSec, SSL VPN, IDS, and IPS. Additional Skills: Routing and Switching: Prior experience in routing and switching. Communication: S
Get new reliability engineer jobs by email
Daily job updates · Unsubscribe anytime