Jobs in United States

Software Reliability Engineer in United States

2,007 active opportunities · Updated October 2026

Explore current software reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
P
📍 New York, NY, United States· Full-time· Hybrid
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Backend Software Engineers at Palantir build software at scale to transform how organisations use data. Our Software Engineers are involved throughout the product lifecycle, from idea generation, design, prototyping, and production delivery. You will collaborate closely with technical and non-technical teammates to understand our customers' problems and build products that solve them. We encourage movement across teams to share context, skills, and experience, so you'll learn about many different technologies and aspects of each product. Engineers work autonomously and make decisions independently, within a community that will support and challenge you as you grow and develop, becoming a strong technical contributor and engineering leader. Your day-to-day workflow will vary, adapting to the requirements of our users and the technical challenges that arise. One day, you may find yourself collaborating with other engineers to architect a new system that enables a novel workflow, the next you could be fine-tuning performance to enable low-latency operational outcomes. Our Product Development organisation is made up of small teams of Software Engineers. Each team focuses on a specific aspect of a product and work collaboratively to build cross functional capabilities, streamline user workflows and continuously improve our software's efficiency and reliability. We’re hiring engineers who are passionate about solving real-world problems and empowering both developers and end-users to work optimally. If you’re motivated to develop reliable, performant, and scalable systems, and to design robust APIs and primitives, this role offers the opportunity to make a

PythonJavaC++Supply Chain
O
📍 Atlanta, Georgia, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

PythonJavaSQLAWS
O
📍 San Francisco, California, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge As a Senior Staff Software Engineer, you will serve as a technical leader for OneTrust’s AI Governance (AIG) platform, driving the design, scalability, and reliability of systems that enable enterprises to deploy and govern AI and LLM-powered applications responsibly. You will deeply understand how customers build, deploy, and operate AI systems, and translate those needs into secure, compliant, and observable platform capabilities. Your Mission Development Lead the design and development of Java/Python microservices and shared libraries integrating with AI platforms for OneTrust’s AI Governance product. Design, build, and test cloud-native applications deployed on Microsoft Azure using Core Java, REST, and the Spring ecosystem. Lead the architecture and development of reusable AIG reporting and dashboard capabilities that integrate governance data from SQL databases and analytical platforms with runtime observability signals. Design reusable semantic-layer and metric-abstraction capabilities, including dataset contracts, metric defini

PythonJavaSQLAWS
M
📍 O Fallon, Missouri, United States
✓ High-confidence listingCompany trend +212.5%
Quick readStrong listing-quality and freshness signals

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Job Overview: Responsible for the analysis, design, development, testing, and delivery of secure, scalable software solutions. Define requirements for new applications and customization adhering to Mastercard standards, processes, and best practices. Develop, customize, and test applications to integrate to Mastercard specifications. Provide leadership, mentoring, and technical training to other team members. Major Accountabilities • Plan, design, architect, and develop secure, scalable, and maintainable technical solutions and alternatives to meet business requirements in adherence with Mastercard standards, processes, and best practices • Lead day-to-day system development and maintenance activities of the team to meet service level agreements (SLAs) and create solutions with a high level of innovation, cost effectiveness, quality, reliability, and faster time to market. • Accountable for the full systems development life cycle including creating high-quality requirements documents, use cases, designs, and other technical artifacts including but not limited to detailed test strategies, performance benchmarking, release rollout and deployment plans, contingency/back-out plans, feasibility studies, cost and time analysis, and detailed estimates. • Design, develop, test, dep

JavaDockerGitAI
C-
📍 New York, NY, United States
✓ High-confidence listing

$325K – $500K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We're looking for a Staff Software Engineer to help build the next generation of CLEAR's enterprise identity platform, CLEAR1. We’re creating frictionless identity solutions for some of the most sophisticated companies in the world and pushing the boundaries everywhere we go. As a Staff Engineer, you’ll own high-impact, cross-functional products and integrations end to end from problem framing and architecture through implementation, rollout, adoption, and measurable outcomes. This is an ideal role for someone who thrives in scrappy, ambiguous environments, while bringing the technical rigor and operating discipline developed at larger companies. You’ll partner across teams, establish reusable patterns and standards, improve reliability and execution, influence technical direction, and mentor other engineers. A brief highlight of our tech stack: Python / Java / React / Typescript What you’ll do: Design, build, test, and deploy scalable applications that power CLEAR’s identity platform. Own projects end to end—from technical discovery and architecture through implementation, rollout, adoption, and operational support. Partner closely with Product, Design, Data, Security, and Operations to translate business problems into simple, scalable technical solutions. Design for reality by understanding failure modes, system dependencies, degradation strategies, recovery paths, and the points where systems may break under scale. Build quality and operability into the design from the start, including test strategy, observability, alerting, sa

TypeScriptPythonJavaReact
D
📍 New York, New York, United States
✓ High-confidence listingCompany trend -85.2%
Quick readStrong listing-quality and freshness signals

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at a high scale - trillions of data points per day — providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity We are looking for an experienced software engineer to join our CI/CD Security Team within our SDLC Security organization. We work at the intersection of security and engineering infrastructure to secure Datadog's continuous integration and continuous delivery systems. Our responsibilities include hardening pipelines, protecting credentials, and enforcing tightly scoped access controls. We also develop authorization and verification mechanisms to ensure that only trusted code and approved processes can reach production. In this role, you will shape and build a new security layer for our CI/CD infrastructure and drive its adoption across the engineering organization. You will solve challenging systems problems around trusted build provenance, secure secret delivery, and real-time policy enforcement at high throughput. The work sits directly in the critical path of software delivery, where strong security guarantees have to coexist with low latency, high reliability, and a seamless developer experience. You’ll join at an ideal time to make a big impact, as the need for robust software supply chain security is higher than ever. Datadog is growing rapidly, and AI-assisted development is increasing both the pace of software delivery and the amount of activity flowing through our CI/CD systems. Securing that scale without slowing engineers down requires strong software engineering fundamentals, thoughtful automation, and security controls designed to operate reliably at high throughput. At Datadog, we pla

JavaScriptPythonJavaAWS
RH
📍 Denver, Colorado, United States
✓ High-confidence listing

$118.7K – $160K/yr

Quick readStrong listing-quality and freshness signals

Healthcare is complex. We’re here to change that. RVO Health is a health technology company on a mission to make health easier to navigate, more accessible, and more affordable for everyone. Here, you'll help over 40 million people every month, with a team that genuinely cares about the work and each other. AT A GLANCE The Senior Software Engineer is a crucial role within our organization, requiring work in various capacities and adaptation to different work arrangements based on the needs set by the business. The successful candidate will be responsible for fulfilling their job duties in the following work situations: Where You'll Be Location: Denver, CO | Hybrid We believe great collaboration happens when we're together, solving problems, learning from each other, and connecting as a team. That's why we’re in our offices Tuesday through Thursday each week. You are welcome to work remotely Mondays and Fridays if you wish. Address: 1801 California St. Denver, CO 80202 What You’ll Do Lead the end-to-end design, development, and implementation of sophisticated software applications and systems aligned with business goals. Collaborate closely with stakeholders including product managers, designers, and other engineers to gather requirements and translate them into robust technical designs and solutions. Write high-quality, efficient, maintainable, and scalable code adhering to best practices and company standards. Debug, analyze, and resolve complex software defects and performance bottlenecks to ensure optimal system reliability and user experience. Conduct comprehensive testing and validation including unit, integration, and performance testing to guarantee software quality. Mentor and provide technical guidance to junior and mid-level engineers, fostering professional growth and knowledge sharing. Perform thorough code reviews to maintain high code quality, enforce coding standards, and promote best p

TypeScriptPythonReactNode.js
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready for their most important workloads. In this role, you help design and run services that sit at the intersection of hardware, security, and large-scale distributed systems. We partner closely with security, silicon, platform, and cloud teams to bring new ideas into reliable production services that people rely on every day. We care about building systems that last, supporting each other, and creating space for learning and experimentation. If you enjoy solving complex problems, keeping services running smoothly, and collaborating with teammates from many disciplines, we would love to talk with you! What you’ll be doing: Your main focus will be on building and managing our core attestation cloud services. Day-to-day responsibilities include crafting APIs and integrations, boosting reliability, and working alongside NVIDIA teams to convert hardware trust mechanisms and standards into production-ready solutions. You will contribute significantly to shaping how customers verify that NVIDIA platforms are secure and prepared for their workloads. Crafting and evolving attestation cloud services, APIs, and SDK/CLI integration points that confirm the integrity of NVIDIA platforms across data center, AI, networking, and partner environments. Improving reliability and operational maturity through SLOs/SLIs, alerting, runbooks, incident response, and safe rollout practices. Crafting resilient service behavior that handles dependency failures, caching challenges, regional issues, customer-side resilience needs, and graceful degradation. Architecting trust-material distribution for certificate status, re

JavaAWSAzureGCP
M
📍 Minnesota, United States of America, United States
✓ High-confidence listingCompany trend +1850%
Quick readStrong listing-quality and freshness signals

We anticipate the application window for this opening will close on - 28 Sep 2026 Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72+ million patients each year. Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life At Medtronic, we push the limits of what technology can do to make tomorrow better than yesterday and that makes it an exciting and rewarding place to work. Medtronic Cardiac Rhythm Management (CRM) Patient Care Systems (PCS) organization develops the next generation medical technologies that alleviate pain, restore health, and extend life for millions of patients across the world. Join the PCS Software data services leadership team, where our charter is to build and develop high-performing data engineering teams to ensure products and programs are appropriately resourced and ensure continuous improvement in technical capability, process and compliance within our software ecosystem—including Information Systems Shared Services. We seek a highly motivated, accountable software leader who is passionate about improving reliability, performance, scalability, and cost efficiency of data services. This leader will foster a culture of continuous learning, innovation, accountability, and collaboration that advances engineering excellence, supports clinicians, and improves patients’ quality of life. In this role, you will lead the data services software organization. As a Senior Manager, you will own multi-year plans to build and develop high-performing team

AWSAzureGCPRecruitment
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Plugin Ecosystem team builds the platform and product experiences that let people extend ChatGPT and Codex. We work on plugins, skills, connectors, interactive apps, and open standards like the Model Context Protocol (MCP). We make plugins easy to discover, install, and use, ensure they’re invoked at the right time, and help people find new ways to get value from them. We want anyone to be able to turn a useful workflow into a plugin, share it, and have other people use it. A plugin can package instructions and skills with connections to the tools and data it needs. Our work spans creation and publishing, reliable execution across our products, clear permissions and approvals, and the controls admins need to bring plugins to their organizations. We work closely with research to improve plugin quality as models evolve. About the Role We’re looking for product-minded engineers to build the systems behind plugins and improve how models use them. Depending on your focus, you may scale generalist infrastructure and identity-related integrations across products, or improve plugin quality at the intersection of backend engineering and applied AI or work on the product experience itself to drive plugin usage. You’ll work across teams and own problems from diagnosis and design through implementation and release. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and ship APIs, SDKs, and services that developers use to extend ChatGPT and Codex. Build intuitive experiences that help users discover, install, and use plugins to get more done. Make plugins easier to create, test, publish, update, and share. Improve when and how models use plugins, from choosing the right plugin to completing a task. Work with Research to diagnose failures and measure improvements as models evolve. Improve plugin reliability and interaction quality acros

Artificial IntelligenceAI
N
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -87.4%
Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion’s Data Foundations team builds and operates the batch and streaming infrastructure behind our product features, analytics, search, and AI experiences. We’re looking for a hands-on technical leader to shape the next generation of this platform as Notion serves larger customers, expands globally, and supports more data-intensive products. You’ll identify the highest-leverage problems, set direction, build and develop a high-performing team, and lead multi-quarter initiatives across our data lake, streaming, distributed-compute, governance, and reliability systems. You’ll stay close to critical technical decisions while creating clear ownership, growing engineers and technical leaders, and helping the team execute as one—partnering closely with Data Engineering, Data Product, Search, AI, Infrastructure, and Security. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You

P
📍 San Francisco, CA, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.3%

From $1.5M/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Job Title: Software Engineer II, Data Analytics and Engineering Intro: We’re looking for a Software Engineer II, Data Analytics and Engineering to improve the quality, reliability and velocity of data science and product development at Pinterest. You’ll build scalable data foundations, analytics tooling and analysis pipelines that enable trusted, self-service access to datasets, insights and metric investigations across cross-functional teams. What you’ll do: Develop and document practical instrumentation and experimentation standards, then partner with product engineering teams to apply them to priority product development work. Build and improve scalable analysis pipelines and tooling that produce reliable insights at scale and strengthen understanding of key data structures and metrics. Create tools and processes that enable Data Scientists a

PythonSQLAWSRest
RH
📍 Denver, Colorado, United States· Full-time
✓ High-confidence listing

From $118.7K/yr

Quick readStrong listing-quality and freshness signals

Healthcare is complex. We're here to change that. RVO Health is a health technology company on a mission to make health easier to navigate, more accessible, and more affordable for everyone. Here, you'll help over 40 million people every month, with a team that genuinely cares about the work and each other. AT A GLANCE The Senior Software Engineer is a crucial role within our organization, requiring work in various capacities and adaptation to different work arrangements based on the needs set by the business. Where You'll Be Location: Denver, CO | Hybrid We believe great collaboration happens when we're together, solving problems, learning from each other, and connecting as a team. That's why we're in our offices Tuesday through Thursday each week. You are welcome to work remotely Mondays and Fridays if you wish. Office Address: 1801 California St. Denver, CO 80202 What You’ll Do Lead the end-to-end design, development, and implementation of sophisticated software applications and systems aligned with business goals. Collaborate closely with stakeholders including product managers, designers, and other engineers to gather requirements and translate them into robust technical designs and solutions. Write high-quality, efficient, maintainable, and scalable code adhering to best practices and company standards. Debug, analyze, and resolve complex software defects and performance bottlenecks to ensure optimal system reliability and user experience. Conduct comprehensive testing and validation including unit, integration, and performance testing to guarantee software quality. Mentor and provide technical guidance to junior and mid-level engineers, fostering professional growth and knowledge sharing. Perform thorough code reviews to maintain high code quality, enforce coding standards, and promote best practices across the team. Continuously improve software development processes, tools, and methodol

JavaScriptTypeScriptPythonJava
B
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$240K – $285K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Staff Software Engineer in Banking, you will help shape the technical direction of one of Brex’s most strategic and complex product areas. The Banking org is both a product and platform org, it owns the Brex Business Account product and AP offerings like Bill Pay and Vendors, and the underlying money movement platform and partner integrations that power those experiences. In this role, you’ll work across customer-facing product surfaces and core financial infrastructure, driving architecture, reliability, and

AIExcelAccountingFinance
🔔

Get new software reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime