About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for an exceptionally experienced, hands-on full-stack engineer to define and build the next generation of AI-powered enterprise workflows. You will take on the hardest and most ambiguous problems in the Technology vertical: translating real customer needs into product direction, designing the systems behind the experience, and personally writing and shipping production-quality code across the stack. You will own the technical direction and end-to-end delivery of products spanning ChatGPT Work surfaces, backend services, plugins, connectors, enterprise data, permissions, and evaluations. You will make foundational architecture and product tradeoffs; establish patterns other engineers can build on; and hold these experiences to a high bar for reliability, security, observability, and customer value. This is an individual-contributor role for an engineer who leads through technical judgment, direct execution, and influence—not people management. You should be equally comfortable working directly with customers, setting direction with senior cro
Jobiba hiring network
Senior Software Reliability Engineer Jobs
7,292 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior software reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic patterns, complex dependency graphs, and frequent change. The ChatGPT Infrastructure team builds and operates the platforms that enable fast iteration without compromising performance or reliability. We design shared systems, data paths, rollout mechanisms, and reliability guardrails that teams rely on to ship changes to ChatGPT at scale. We focus on high-leverage infrastructure: primitives and “golden paths” that incorporate operational lessons as defaults, so engineers don’t need to rediscover failure modes, latency pitfalls, or integration issues each time they build something new. About the Role We’re hiring Senior and Staff Engineers to design and build infrastructure systems that underlie ChatGPT and multiply the effectiveness of teams building user experiences. This is not a support-only role. It’s a platform-building role: you’ll define interfaces, develop core abstractions, and create tooling to make safe, fast iteration the norm. Your work will reduce friction, prevent regressions, improve performance, and ensure systems scale gracefully as the product grows. Where You Can Have Impact You might work on one or more of the following areas (without being restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution. Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases. Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression contr
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is expanding the boundaries of the Data Cloud to support mission-critical transactional workloads. Our goal is to deliver OLTP capabilities with the performance, reliability, simplicity, and scale customers expect from Snowflake, while creating a seamless experience across transactional and analytical data. We are looking for a Senior Engineering Manager – OLTP to lead engineering teams and leaders building core transactional database technology and the cloud infrastructure required to operate it at scale. You will help define the architecture and roadmap, grow the organization, and drive technology from design through production. AS A SENIOR ENGINEERING MANAGER – OLTP AT SNOWFLAKE, YOU WILL: Set technical and execution strategy for key areas of Snowflake's OLTP platform, translating product goals into architecture, roadmaps, and team plans. Lead and grow multiple engineering teams, developing managers and senior technical leaders while fostering a culture of ownership, technical excellence, and execution. Drive adoption of AI and agentic development practices to improve engineering velocity, quality, and productivity across the software development lifecycle. Drive architecture and technical decisions in areas such as transactions, concurrency control, low-latenc
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - Container Platform (Menlo Park) About the Role We build the foundational container platform that runs Snowflake's production, AI/ML, and CI workloads across AWS, Azure, and GCP, including a rapidly growing AI/ML footprint. Hundreds of large Kubernetes clusters under management and growing. The work is to make that fleet reliable, automated, and invisible to the thousands of engineers building on top of it. This is a staff-level role on a senior, high-performing platform team. You'll own hard problems end to end, drive technical direction across teams, and build the automation and platform abstractions that make operating at this scale sustainable. There is significant unsolved work ahead: improving the developer experience for thousands of internal engineers and continuing to scale the platform to meet Snowflake's growth. What You'll Do Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling. Identify and drive cross-team technical improvements across the platform, from architecture through adoption. Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality. Act as a technical anchor for the team, dev
About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last-mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We're looking for an experienced technical operator to lead live testing, deployment, and operational validation of cutting-edge autonomous technologies. This role sits at the intersection of engineering and operations, helping ensure new capabilities are safely deployed, thoroughly evaluated, and translated into actionable engineering feedback. You’re excited about this opportunity because you will… Lead and oversee live testing across transport, deployment, and safety validation. Partner closely with hardware, software, and business operations teams to validate new product capabilities while providing guidance and mentorship to junior team members. Conduct and document complex tests for autonomous technologies, evaluating robot behavior, identifying issues, and validating new features and requirements. Provide actionable technical feedback to engineering teams based on test outcomes. Exercise technical judgment during live testing by evaluating robot behavior, assessing operational risk, distinguishing expected behavior from product defects, and determining when engineering escalation or additional validation is required. Develop and implement testing processes, protocols, and checklists that improve the safety, efficiency, and reliability of new products. Utilize internal tools to analyze logs, investigate issues, document findings, and track issues through resolution. Mentor junior team members in structured debugging and documentation practices. Conduct detailed analyses and generate comprehensive reports that identify trends, summarize findings, and provide strategic recommendations to engineering and operations partners. We’re excited about you because… 2+ years of exper
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview The Shopper Activation & Engagement team at Instacart owns core moments in the shopper lifecycle, including onboarding, activation, finding work, evaluating and accepting batches, staying compliant, understanding earnings-related experiences, and remaining engaged on the platform. The team's work directly supports shopper experience, supply quality, marketplace efficiency, operational reliability, and compliance. We are looking for a Senior Manager, Software Engineering to lead a team of 10–12 engineers across mobile, backend, vendor-integrated systems, and engineering foundations. This leader will shape the team's strategy, execution model, technical direction, and talent development while partnering closely with Product, Design, Data, Operations, Legal/Compliance, and other cross-functional teams. This is a high-impact role for someone who thrives in
NVIDIA has been redefining computer graphics, desktop gaming, and enhanced computing capabilities for more than 25 years. Today, we are tapping into the unlimited potential of AI to define the next era of computing. As a NVIDIAN, you will work on problems that sit at the boundary of architecture, silicon, firmware, software, and production, where strong judgment matters as much as technical depth. We're the Silicon Design for Productization (DFP) Team, within the broader Silicon Co-Design Group, and we turn power and thermal design into executable productization methodology. Power and thermal are among the most complicated problems we work on at NVIDIA because they sit at the intersection of architecture, workload behavior, silicon variation, firmware policy, platform constraints, and product goals. Small decisions here have an outsized impact on performance, efficiency, reliability, bring-up speed, and ultimately what the product can deliver in the field. We define how features move from concepts to bring-up, characterization, validation, and release. In this role, you will help us build that bridge. We're looking for an engineer who reasons from first principles, flourishes with ownership in a fast-paced environment, and uses AI with sound judgment. What you’ll be doing: Lead the effort across multi-functional teams to keep the program’s power and thermal productization strategy clear, executable, and on track. Create methodology and silicon test plan based controller designs and architecture, including characterization process, debug tools, fuse/firmware settings and lab requirements. Drive resolution for challenging silicon issues through structured hypotheses, measurement plans, and root-cause closure. Steward the Power and Thermal playbook when the existing productization methodology
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Driver team is dedicated to fostering a platform of high-quality service by empowering drivers to perform their best. We are looking for a product-minded engineer who wants to build and improve products that sit at the center of the core driver experience. Products you drive will solve pain points that matter most to drivers by streamlining key interactions, reducing friction, and creating systems that feel intuitive, fair, and supportive. As an engineer at Lyft, you'll collaborate with teams like product, data science, analytics, and operations on code that empower us to iterate quickly, while focusing on delighting our passengers and drivers. Responsibilities: Design and implement backend features end-to-end with clear ownership, delivering well-scoped work from technical design through to production with moderate guidance from senior engineers Write clean, reliable, well-tested code that meets team standards and holds up in code review Participate actively in code reviews, giving specific and constructive feedback while continuing to develop your own review instincts Debug and resolve issues across backend services including performance bottlenecks, reliability problems, and data integrity issues Collaborate with product managers, designers, and partner engineering teams to clarify requirements and surface technical constraints early Contribute to technical discussions and help evaluate implementation approaches for new features Write unit and integration tests for your own code and develop familiarity with the team's broader testing and observability practices Address technical debt and make incremental improvements to existing services as part of regular development work Participate in on-call rotations, respond to production incidents, and support teammates in mitigating custome
Position Overview We are looking for a Software Engineer II to build and deliver scalable software solutions across our products. You will work on modern web applications and cloud-based services using Node.js, React, TypeScript, AWS, PostgreSQL, MSSQL, and Docker, while contributing to AI-enabled features and integrations. You will collaborate closely with other engineers, product managers, and cross-functional teams to develop reliable, maintainable, and production-ready solutions. This role provides an opportunity to work with modern AI technologies including Python, AWS Bedrock, MCP, RAG, and agentic AI workflows while developing strong expertise in cloud-native software engineering. What You'll Do Develop and maintain scalable backend services and APIs using Node.js, TypeScript, and JavaScript. Build responsive and maintainable frontend applications using React. Design and implement integrations with AWS services and contribute to cloud-native application development. Develop and maintain applications using PostgreSQL and MSSQL, including writing efficient queries and working with database schemas. Build, test, and deploy applications using Docker and modern CI/CD practices. Contribute to AI-enabled product features using Python, AWS Bedrock, RAG, MCP, and AI integration patterns. Work with the team to integrate LLM capabilities, APIs, tools, and data sources into production applications. Write clean, maintainable, and well-tested code following established engineering practices. Participate in code reviews, technical discussions, debugging, and production issue resolution. Develop unit and integration tests and contribute to improving application quality and reliability. Monitor application performance and troubleshoot issues across development and production environments. Collaborate with senior engineers and architects to implement technical solutions aligned with product and engineering requirements. Stay current with emerging technologies, particularly in
Lithic is the modern card issuing and processing platform empowering ambitious financial companies to build the future of payments. Our infrastructure powers card programs for 100+ innovative clients, from fintechs reimagining credit and digital banking to platforms transforming disbursements and spend management. Companies like Mercury, Flex, and Novo rely on Lithic's developer-friendly APIs, direct network connections, and flawless reconciliation to launch and scale card programs in weeks, not years. We're building a future where access to better financial products materially improves people's lives, free from the constraints of 30-year-old mainframes and legacy processors. We're proud to be backed by world-class investors who share that vision, including Bessemer Venture Partners, Index Ventures, Spark Capital, Stripes, and Mastercard, along with many others. We're a team of 170+ across 26 states and 7 countries, headquartered in New York City. We are hiring for our Treasury team Software Engineers at various levels (II and Senior) who are curious and willing to dive deep and understand our technology and domain in order to solve interesting and hard problems.The Treasury team maintains and builds the backend services that manage the flow of funds between Lithic and third parties. This includes our ledger, ACH and wire infrastructure, and associated reconciliation. The systems we maintain have high standards of reliability and correctness. You will become an expert in the card payments space. The Treasury team primarily uses Python for their tech stack. What You'll Do: Ensure high reliability and correctness for Lithic’s ledger and orchestrated funds flows Develop new features to better serve Lithic customers Ensure that the team is delivering reliable, secure, and scalable code with minimal tech debt Own initiatives from planning to launch, keeping stakeholders informed and aligned along the way Lead efforts to improve systems and processes wit
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we are building the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a individual contributor role for a hands-on system engineer who executes moderately complex tasks, builds platform components and collaborates under senior guidance. What You'll do : Contribute technical design and execution for automation initiatives within the team, creating paved paths that enable builders across Okta to connect enterprise systems seamlessly. Design, build, and deploy high-throughput event-driven integration flows , API gateways, and async orchestration workflows using AWS architectures and iPaaS platforms. Build reusable frameworks , developer SDKs, self-service primitives, and integration templates to streamline automation delivery. Develop and maintain AWS cloud services and modern iPaaS tooling, driving scalable architecture, reliability, and builder enablement. Partner with operations, security, and platform teams to strengthen monitoring, observability, structured audit logging, and automated governance. Actively participate in code reviews, team agile ceremonies, and technical discussions Collaborate with security, compliance, and business stakeholders to ensure self-service automations are secure, resilient, zero-tr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. This is a contract position through our staffing partner Magnit. 108,000.00 - 135,000.00 - 162,000.00 CAD Annual This role is not eligible for the Okta-sponsored benefits listed below. Magnit will provide any locally required benefits. Okta seeks a skilled Senior Recruiter to drive full-lifecycle recruitment and build strategic talent pipelines for our Engineering organization across North America. As part of our AMER Tech Recruiting team, you will be a trusted talent advisor responsible for sourcing, engaging, and delivering top-tier engineering talent while maintaining an "always recruiting" mindset in a fast-paced, high-growth environment. What You'll Be Doing Own full-lifecycle recruitment for Engineering and technical roles (Software Engineering, Site Reliability, Security, TPM, Product) across US & Canada, managing a flexible req load that scales with business priorities. Partner strategically with hiring managers and leadership to understand talent needs, define role scope, advise on talent gap mitigation, and challenge assumptions to ensure hiring decisions strengthen long-term organizational capability. Build and execute talent strategies that balance external hiring with internal mobility, creating sustainable pipelines that reflect commitment to diversity, inclusion, and high-performing engineering culture. Drive metrics-informed recruiting decisions by developing KPIs, analyzing recruiting data, and using insights to optimi
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring talented Software Engineers for the Snowflake Dynamic Tables team in Berlin, Germany. Join us to build the next generation data platform that enables customers to transform data with declarative SQL while maintaining control over cost, latency, and throughput. We are looking for strong engineers who are enthusiastic about building new cutting-edge technologies, who look forward to tackling complex database problems, and pick up and understand deep technical areas quickly. You will work alongside seasoned engineers and grow in your scope and influence. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Work with a talented and collaborative team of engineers and Product Managers in our globally distributed team to design and build Dynamic Tables capabilities Design, implement, support, and evolve new features and performance improvements Help shape technical and product direction with senior teammates Break ambiguous problems down, weigh the tradeoffs, and make technical and product decisions Analyze and solve performance, correctness, and fault-tolerance challenges at scale Dig into unfamiliar parts of a large system to root cause and solve problems Ensure operational readiness of what you build and help meet the commitments to our customers regarding reliability, a
ABOUT THE ROLE Mid level Software Engineer will be a member of the agile team which is responsible for design and development of scalable microservices for Peloton's core features across all platforms (Bike, Tread, Strength, and Digital). The Content AI Team is responsible for training and hosting various models in the domains of Natural Language Processing (NLP) and Automatic Speech Recognition (ASR). The candidate will be responsible for developing, testing, deploying, and monitoring microservices that specifically power features like search, voice, and subtitles, which are crucial for platform expansion and international growth. In addition to technical delivery, good communication skills are essential for providing updates to engineering leads and other stakeholders, such as Technical Program Managers (TPM). YOUR DAILY IMPACT AT PELOTON Develop and maintain business-critical APIs and services with a focus on high availability, low latency, security and scalability under guidance from senior engineers Effectively provide updates to the team leaders and participate in sprint planning and team meetings Write understandable, well-tested code with an eye towards maintainability and scalability Contribute to building reusable code, libraries, and patterns for use across teams Implement solutions to scale services while meeting business and product requirements, often under supervision Utilize production monitoring/profiling/tracing and load testing tools to discover bottlenecks and apply techniques such as data modeling, query optimization, and caching to address them Understand and implement industry best practices such as feature toggles, CI/CD, test automation, logging, and monitoring in order to ensure confidence in our release process Participate in on-call rotations and assist with incident response efforts to maintain system reliability YOU BRING TO PELOTON MS or BS degree in fields such as Computer Science, Engineering or Mathematics 3+ years of software devel
Get new senior software reliability engineer jobs by email
Daily job updates · Unsubscribe anytime