At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Where Data Does More. Join the Snowflake team. Join our ML Feature Store team where we're building cutting-edge product capabilities that power complex feature transformations and low latency feature serving. We're revolutionizing machine learning feature management and serving capabilities as part of the Snowflake ML suite of products. In the era of GenAI and agents, our team delivers high-quality, fresh feature solutions that make a real difference for our customers. IN THIS ROLE AT SNOWFLAKE, YOU WILL: Help define and own the roadmap for Snowflake Feature Store, working collaboratively with senior architects and ML team leadership Build and execute a vision for incorporating new advances in machine learning Ensure operational excellence of services and meet reliability, availability, and performance commitments Collaborate across ML partner teams to improve development velocity and capabilities Support team members in delivering high technical quality WE WOULD LOVE TO HEAR FROM YOU IF YOU HAVE: 10+ years of experience in designing and building data serving infrastructure and/or machine learning platforms. Strong track record working with machine learning systems and platforms. Strong understanding of computer science fundamentals. B.Sc . in Computer Science Fluency in Ja
Jobs in United States
Software Reliability Engineer in United States
2,007 active opportunities · Updated October 2026
Showing
15 jobs
Explore current software reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Astral team builds high-performance developer tools to power the future of programming, at OpenAI and beyond, including Ruff, uv, and ty. The Astral toolchain sees hundreds of millions of installs per month and powers hundreds of millions of package downloads per day for the Python ecosystem. As a team, we are building on those foundations to continue solving impactful tooling problems as programming evolves. About the Role We are looking for an experienced software engineer to build next-generation programming language tooling. If you like writing high-performance Rust, it could be a good fit; if you like thinking about the future of programming, it could also be a good fit. Strong candidates tend to have deep experience with Rust, Python, open source, compilers, or developer tools — but few candidates are deep in all of these areas, and we've hired candidates without prior Rust or Python experience. In this role, you will: Design and implement features in Astral’s existing open source projects (Ruff, uv, ty, and python-build-standalone, and more). Support Astral’s open source projects as a maintainer, triaging user issues, reviewing pull requests, and participating in community discussions. Evolve the Astral toolchain to accelerate development velocity at OpenAI. Build entirely new tools, in entirely different programming ecosystems, to power the future of agentic software development. Your background might look something like: 5+ years of professional engineering experience, excluding internships, in relevant engineering roles. High agency and comfort operating in a fast-moving environment, with strong ownership of security, reliability, and operational excellence. Strong developer empathy and communication skills, including experience maintaining open source projects. Exceptional systems engineering fundamentals and a track record of leading complex projects from ambiguous problem statements through to user impact. Proficiency in one or more s
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. As a Senior Software Engineer, you'll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, driving complex projects across our technical stack, and mentoring our talented engineering team. Your past experience will be leveraged to enable and accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Authoring Experience team owns how Vanta tests are authored and deployed to customers, from the APIs and services underneath to the experiences built on top. As a Senior Backend Engineer on this team, you will own the services and public-facing customer APIs that turn Vanta's automation platform into products customers touch directly, stitching together systems across the platform to serve them well. What you’ll do as a Senior Fullstack Software Engineer at Vanta: Own the architecture and rollout strategy for critical platform services, from initial design through large-scale rollout and migration. Set the technical bar for how we design, publish, and maintain customer-facing APIs while owning things like security, reliability, and developer experience. Stitch together
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team responsible for building the authentication and security solutions that help protect digital interactions across Mastercard's Identity Solutions platform. As a Senior Software Engineer, you will design, develop, and support secure, scalable, and high-performing applications that enable trusted identity verification, authentication, and fraud prevention capabilities. In this role, you will take ownership of complex technical challenges, contribute to software design and architecture decisions, and partner closely with product, security, and platform teams to deliver reliable, production-ready solutions. You'll play a key role in advancing engineering excellence through secure development practices, system reliability, automation, and continuous improvement while mentoring other engineers and influencing technical direction across the team. Role •Design, build, test, deploy, and maintain scalable, cloud-native applications and microservices •Develop REST APIs using Java and Spring Boot, focusing on performance, scalability, and reliability •Translate requirements into well-structured designs and architecture, ensuring maintainability and security •Lead and contribute to system design discussions, aligning with architectural standards and best practices<
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Software Engineer II Overview Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. The Fraud Products team (part of O&T) is developing new capabilities for MasterCard's Decision Management Platform, which serves as the core for multiple business solutions to combat fraud and validate cardholder identity. Our patented Java-based platform processes billions of transactions per month in tens of milliseconds using a multi-tiered, message-oriented approach for high performance and availability. MasterCard software engineering teams leverage Agile development principles, advanced development, design and test automation practices, and an obsession over security, reliability, and perfo
Job Requisition ID # 26WD100611 Position Overview Autodesk is a global leader in design and make technology, with expertise across architecture, engineering, construction, design, manufacturing, and entertainment. Our software and services empower innovators everywhere to solve challenges big and small—from greener buildings to smarter products to more compelling media and entertainment. At Autodesk, we believe that when you have the right tools to work and think flexibly, you have the power to transform what actually needs making. We provide our customers with technology to help them achieve better outcomes for their products, businesses, and the world. We are looking for a highly motivated software engineer to join the PSET-Access group. The team builds and operates systems that enable secure, reliable, and scalable access to product capabilities and resources, and now in the process to develop the next generation to modernize the access management capabilities. In this role, you will work with engineers and cross-functional partners to design, build, test, deploy, and operate production software. You will primarily develop backend services using Java or Go, while owning well-defined components and projects, contributing to technical decisions, and helping improve the reliability, maintainability, and developer experience of the systems the team owns. This is a strong opportunity for an engineer who has solid software engineering fundamentals and is ready to grow their technical depth and ownership. Responsibilities Design, implement, test, deploy, and maintain production-quality backend software, prim
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will work on the systems software strategy and execution that brings new AI silicon from first power-on to a fully integrated system running production-representative models at expected functionality and performance. You will define how software exercises and validates compute, memory, interconnect, and I/O subsystems, then build the diagnostics, automation, and observability needed to find issues quickly. This role sits at the center of silicon, firmware, platform, systems, and workload teams. You will turn hardware specifications and performance targets into an end-to-end bringup plan, drive cross-functional debug, and establish the stress and regression infrastructure that makes each new platform reliable across operating environments. In this role, you will: Contribute to the end-to-end software bringup and validation strategy for new silicon and first-party systems. Define software-driven test coverage across compute, memory, interconnect, I/O, and their system-level interactions. Build diagnostics, test automation, telemetry, and regression infrastructure that accelerate first-silicon learning and issue isolation. Lead bringup from initial silicon arrival through board and system integration, docking, runtime enablement, and model execution. Design stress tests that characterize reliability, performance, and stability across workloads and operating conditions. Translate architecture specifications and performance models into measurable acceptance crit
$213K – $320K/yr
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: As a Developer Platform engineer, you will directly contribute to the foundational pieces that make Notion extensible and connected. You will build tools, APIs, and platform experiences that help customers connect Notion to the world — bringing their apps, data, and workflows into one workspace. Your work will make it easier for developers, admins, and builders to create reliable integrations and automations on top of Notion. You will be a key player in building the robust technical foundation that allows Notion to achieve the connected workspace vision. Your work will include both internal platform contributions that accelerate other Notion engineering teams and end-user-facing functionality that enables toolmaking ubiquity. You will be presented with challenging technical problems, as Notion’s product needs are complex. You’ll play a key role in identifying and executing against technical investments that ensure the long-term quality, reliability, and performance of Notion’s platform as we scale. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursd
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re searching for a highly motivated, technical leader to design, drive, and operationalize rack-scale factory and deployment flows for next-generation data center products. The ideal candidate will combine deep systems expertise, decisive technical leadership, and a passion for building reliable, debuggable, and scalable manufacturing and deployment solutions. What you’ll be doing: Lead and drive rack-scale/L11 flows for factory and initial data center deployment. Design and implement end-to-end factory workflows, including firmware flashing sequences, security provisioning, and deployment of software mitigations. Collaborate with data center architects, ODMs, and OEMs to define factory and data center requirements that ensure efficient and reliable production ramp. Champion reliability, debuggability an
About the Team OpenAI's mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. The API Platform turns frontier research into reliable capabilities that developers use to build transformative products and services for people around the world. API Safety's goal is to ensure safe deployment of frontier models in the API. We design APIs and systems that help developers share usage context, understand safety events, and apply safeguards tailored to the risk profile of the applications they are building. This work is critical to our frontier model launches and partners closely with teams across API, Integrity, and Safety Research. About the Role We're looking for product-minded software engineers to join a team that is addressing emerging risks at the frontier of model development while building novel solutions for real-world AI deployment. The day-to-day work ranges from solving production challenges to designing new product experiences and safeguards. The right candidate is comfortable balancing tradeoffs across developer experience, latency, reliability, and risk. In this role, you will: Design and build dashboards and APIs for safety controls and customer-facing observability. Develop scalable systems that extend trusted safety capabilities to new use cases, customers, and deployment environments. Partner with Safety Research and Integrity to build safeguards that mitigate emerging risks. Be responsible for the availability, latency, and scalability of safeguards across high-volume API traffic. Own projects from technical design and implementation through launch and ongoing iteration, while raising the team’s engineering standards Your background might look something like: 7+ years of professional experience, excluding internships, in backend, infrastructure, platform, or product engineering roles. A track record of designing, building, and operating production backend services, developer-facing APIs, or distributed systems. Strong s
As a Senior Backend Engineer on Coder's Enterprise Experience team, you'll build the systems that help large organizations run Coder in production with confidence. You'll improve how Coder scales, how it's upgraded, and how reliably it performs in regulated, air-gapped, and enterprise environments. You'll work on a genuinely cross-functional team of backend, platform, and QA engineers who own the end-to-end experience for Coder operators. From designing new features to evolving Coder's architecture, you'll partner across engineering and product to solve complex problems and ship software that operators trust. What you'll do here Design and build new features end to end, from technical design through production rollout. Design and implement backend architecture changes that support Coder's long-term scalability goals. Investigate and resolve scalability bottlenecks under production-like load, from database access patterns to concurrency handling in coderd. Improve database migration safety and upgrade reliability through schema compatibility, background migrations, and safe rollback strategies. Own the backend side of issues surfaced by Coder operators and administrators. Document the design, implementation, and operational tradeoffs of the systems you build. Participate in code reviews, RFC-style design discussions, and on-call rotations for the services you own. What we're looking for 5+ years of professional software engineering experience, including significant production experience with Go. Deep understanding of Go's concurrency model, including goroutines, channels, the sync package, and debugging race conditions under real-world load. Experience designing and operating relational databases in production, including schema design, migrations, and transactions. Strong verbal and written communication skills. Exceptional debugging and troubleshooting skills, with the persistence to drive complex problems to resolution. A self-motivated, analytical engineer who enj
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for Observe by Snowflake on the Data Management team. This team is responsible for the tables, views, and materialized views at the core of Observe's architecture. Observe's data lake approach lets customers correlate heterogeneous telemetry — logs, metrics, traces, events — across a unified data model. This role owns that data model: how customers define, shape, and query the semi-structured data that makes cross-si
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Senior Software Engineer to own key components of our AI native External Observability Platform . In this role, you will contribute to the technical road map for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex technical projects, and skilled at collaborating with the brightest technical minds in the industry. Key Responsibilities Develop and Scale Distributed Infrastructure: Design and implement key components of Snowf
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an Infrastructure Software Engineer at Baseten, you'll build and maintain components of our ML inference platform that powers production AI applications. You'll contribute to the core infrastructure, enabling developers to deploy, scale, and monitor ML models with high performance. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Infrastructure team: Multi-cloud capacity management Inference on B200 GPUs Multi-node inference Fractional H100 GPUs for efficient model serving RESPONSIBILITIES Develop infrastructure components for our ML inference platform using Python and Go Implement and maintain Kubernetes deployments for model serving Contribute to our inference orchestration layer for model deployments Build and enhance monitoring systems for model performance metrics Implement efficient resource management solutions for ML workloads Support infrastructure automation to improve ML deployment workflows Work closely with team members to implement technical solutions Help balance performance optimization with system reliability Participate in technical discussions around infrastructure improvements Learn and apply infrastructure best practices REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Go proficiency is a plus Working knowledge of Kubernetes and containeriza
Other cities to consider
More places hiring for this role
Get new software reliability engineer jobs in United States by email
Daily job updates · Unsubscribe anytime