Jobs in United States

Reliability Engineer Iii in United States

655 active opportunities · Updated October 2026

Explore current reliability engineer iii jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 New York City, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp's Dev API team builds the programmatic surfaces that let developers and AI agents read from and write to Ramp. We are an AI-focused platform team making it trivially easy for any engineer at Ramp (or outside it) to ship new capabilities across all surfaces. Our ideal candidate thinks agent-first, has strong opinions on developer experience, and wants to define how software interacts with financial infrastructure at scale. Check out our Engineering Blog for more on our tech stack, mission, and values! What You'll Do Build and operate Ramp's multi-surface platform with a high bar for reliability, correctness, and developer experience across tens of thousands of businesses. Serve thousands of builders and partners using our API, driving a large fraction of Ramp’s revenue, and deliver on a massive business opportunity to embed Ramp everywhere. Build software factories -- autonomous tooling and agents that accelerate endpoint creation, enabling teams across Ramp to ship their own API surfaces. Lead design and execution of complex back

PythonRestAIGo
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Streaming Platform team at Sentry is building the next generation of infrastructure that powers our ingestion pipelines and real-time data processing systems. Our platform ingests, processes, and distributes hundreds of thousands of events per second with low latency and high reliability. We are creating a system that makes it easy for Sentry engineers to deploy and run Streaming Applications at scale by simplifying the complexity of Kafka, scaling consumers automatically, and managing state so product teams can focus on building great experiences for developers. As part of this team, you will work on challenges at the intersection of distributed systems, real-time data processing, and developer experience. You will help us create a self-service streaming platform that improves stability, accelerates time to production, and reduces operational overhead. In this role you will Design, build, and operate components of our Streaming Platform, including Kafka, the streaming runtime, high-level APIs, and developer-facing abstractions. Implement resilient, high-throughput stream processing systems that handle unbounded datasets with strong correctness guarantees (delivery, checkpointing, watermarking, and more). Build scalable automation and control plane for Kafka fleet management and improve efficiency. Partner with product engineers to ensure our abstractions enable fast, reliable, and consistent ingestion pipelines. Improve observability, monitoring, and failover for mission-critical real-time systems. You’ll love this job if you You enjoy working on distributed systems at scale and care about reliability and

PythonJavaSQLAWS
R
📍 New York City, NY, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role As a Design Engineer at Ramp, you sit at the intersection of design and front-end engineering. You bring ambitious ideas to life in real interfaces, prototype directly in code, validate with customers, and ship production experiences with a high bar for craft, performance, and reliability. Design at Ramp is AI first and builder led. Work starts in an LLM to clarify intent and constraints, moves into tools like Claude and Cursor to explore and build, is tested with real customers, and then comes into Figma for systems and polish. AI and self-serve research are default parts of the workflow, not side experiments. We are looking for multiple design engineers who think like a designer, code like a front-end engineer, and treat AI as a real collaborator. What You’ll Do Design and build high-craft product experiences: Own key surfaces end to end, from concept through production. Use layout, interaction, and motion to make complex behavior feel simple and safe. Work AI first with Cursor and Claude Code: Use LLMs, Cursor, and Claude Code as y

JavaScriptTypeScriptJavaReact
L
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. The Data team works across Linear, supporting Product, Engineering, and GTM. We own our data pipelines, warehouse, dashboards, analysis, and integrations with third-party tools. As a small team, we focus on building systems that make data accessible and useful across Linear. We’re looking for someone who wants to help shape how we architect, build, and use data as we grow. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in North America. You can work from anywhere within this region. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you’ll do Work across Product and GTM (Marketing, Sales, Customer Success, and Finance) to turn ambiguous questions and operational needs into useful metrics, models, analyses, and workflows Build and maintain dbt models and pipelines that create trusted views of our product, customers, and business Design clear, maintainable data models and improve the testing, documentation, performance, and reliability of our data stack Build dashboards and self-service reporting in Metabase and Hex, and dig deeper when the answer requires more than a chart Operationalize data through reverse ETL and partner with GTM Engineering on the scoring, segmentation, a

R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Job Summary We are looking for an experienced Growth Infrastructure Engineer to build and maintain the technical backbone that enables scalable growth experiments, high-performance data pipelines, and automated systems that drive user acquisition, engagement, and product iteration. This role sits at the intersection of growth, product, and infrastructure — combining deep technical engineering with experimentation and data-driven optimization. You will collaborate with product, data science, and backend teams to ensure that growth initiatives run smoothly and scale efficiently across systems. Key Responsibilities Growth Infrastructure & Systems Design, implement, and maintain scalable infrastructure that supports growth and experimentation needs. Build and optimize analytics pipelines to capture key product and growth metrics (acquisition, activation, retention, etc.). Develop automated workflows for user onboarding, campaign delivery, and performance tracking. Experimentation & Optimization Support A/B testing frameworks and integrate them into production systems. Enable reliable data collection and evaluation for growth experiments. Automate deployment and rollout of growth feature flags and tests. Cross-Functional Collaboration Partner with Growth Product Managers, Data Engineers, and Analysts to define technical requirements for growth initiatives. Translate business goals into technical specifications and system designs. Provide guidance on performance, reliability, and scalability trade-offs. Monitoring & Reliability Implement monitoring and alerting for growth infrastructure services. Troubleshoot production issues and optimize for uptime and performance. Ensure data quality and consistency for report

JavaScriptPythonJavaAWS
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Team Product Platform builds and owns the shared foundations the rest of Replit is built on, spanning the full stack so every other team can ship features safely and quickly: backend infrastructure, connectors, product primitives, and the frontend platform. Our work is high-leverage and horizontal: when our foundations are solid every other team moves faster, and the role gives you exposure across the whole of engineering. We are a small, collaborative team that values curiosity and clear thinking over pedigree, and we work in the open by bringing each other the problem rather than just the request. We care more about how you reason and build than the route you took to get here. About The Role As a Product Engineer , you can focus on frontend, backend, or full-stack work building the shared systems other teams depend on. The work is guided by a few simple questions: Are our shared systems fast, reliable, and cost-efficient as traffic grows? Are we making product development safe by default, consistent, and faster? Can a builder connect a third-party service once and have it work safely across every app they build? Are user-facing surfaces consistent and fast, with shared primitives teams can build on? Is our codebase easy to navigate, change, and extend, including for AI coding agents? What you’ll do Design reusable primitives and interfaces with clear contracts and documentation that other teams adopt Work directly with product teams to turn their friction into platform improvements Profile and instrument shared systems, then ship the improvements that move latency, cost, and reliability Harden systems against failure and abuse, and make safe defaults the path of least resistance Set technical direction in a

TypeScriptReactNode.jsRest
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, o

GCPLinuxAIGo
M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.

PythonAIGo
NR
📍 Portland, Oregon, United States· Full-time
✓ Quality checkedCompany trend -73.9%

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Data, Identity, & API Platform group at New Relic builds the foundation for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the proprietary telemetry database technology at the core of our business. We own our software from top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally be on-call to respond to system failures; so we prioritize work that keeps the lights on and the pager quiet, in addition to the work that powers all of our new products and streams of data. If the idea of working on systems that process millions of messages per second and handle petabytes of data excites you, then you may be an excellent fit! What you'll do Own the New Relic query language and gateway stack Proactively participate in cross-functional committees to move the query language and gateway forward, ranging from collaborations with AI, Visualizations, and Data Processi

JavaAWSAzureGCP
NR
📍 Portland, Oregon, United States· Full-time
✓ High-confidence listingCompany trend -73.9%

From $98K/yr

Quick readStrong listing-quality and freshness signals

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity The Telemetry Data Platform group at New Relic builds the foundation for all of our products: data ingest, storage, and query. As an engineer working on NRDB, you’ll be contributing directly to the proprietary telemetry database technology at the core of our business. We own our software from top to bottom and are directly responsible for its quality and reliability. Each member of the team shares our pager rotation and will occasionally be on-call to respond to system failures; so we prioritize work that keeps the lights on and the pager quiet, in addition to the work that powers all of our new products and streams of data. If the idea of working on systems that process millions of messages per second and handle exabytes of data excites you, then you may be an excellent fit! What you'll do Develop new features with a focus on optimizing performance and efficiency Collaborate with the team to implement scalable solutions and enhance application performance Identifying and acting on opportunities to improve the reliability of our services This role requires 2+ years of professional experience in distributed SaaS software development. Proficiency in Java programming, expertise with algorithms and data structures, and building high-throughput software following best-practices. Deeper understanding of distributed systems and their core challenges. Experience using the command line to manage, investigate, and fix things when they’re broken. Expe

JavaSQLMySQLMongoDB
SF
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $88.1K/yr

Quick readStrong listing-quality and freshness signals

About Stitch Fix, Inc. Stitch Fix (NASDAQ: SFIX) Stitch Fix is redefining retail by combining human creativity with advanced data science and Generative AI. As we build the future of personalized shopping, we’re equally committed to building yours. We believe in investing in our team as much as our technology. Join us to be a trendsetter in the industry and help us redefine what’s possible for our clients, while we help you reach your full potential. About the Role As a Platform Engineer, you will contribute to building and improving Stitch Fix’s cloud-native infrastructure and internal developer tooling. You’ll work on tools and automation that help product engineers deploy, operate, and debug services more easily, while learning modern platform engineering practices alongside experienced teammates. This role is ideal for engineers who enjoy improving developer experience and want to grow their skills in cloud infrastructure and CI/CD systems. Responsibilities: Contribute to the development and evolution of our internal platform-as-a-service used by application and service developers Build and maintain tooling that improves developer workflows, deployment reliability, and day-to-day productivity Collaborate with platform and application engineers to identify friction points and implement incremental improvements Learn and apply best practices around Infrastructure-as-Code, containerized workloads, and CI/CD pipelines Use, or are eager to adopt, AI-assisted development tools to improve productivity, and are excited to help explore and integrate LLM-powered solutions that automate internal support and operational workflows Have opportunities to propose ideas and improvements, with support and mentorship from the team Things you’ll get exposure to (and we don’t expect experience with everything): AWS Terraform, Pulumi CircleCI Docker, ECS, EKS Ruby, Golang, Python About You 2+ years of software development and infrastructure experience with significant contribut

PythonAWSDockerCI/CD
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $153.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We're looking for an outstanding engineer to join the Roblox Datasets team. We develop and maintain highly leveraged core datasets, frameworks, and tooling that support the growing demand for analytics across the company. Our team sits within Foundation AI, and we partner closely with Product, Engineering, Analytics, and Data Science. We help these teams move faster by making core data more reliable, better modeled, easier to discover, and easier to use. This team sits at an important intersection between platform and product. We partner broadly across major Roblox pillars, including Infra, Economy, Creator, Engine, Ads, Discovery, Growth, Compliance, Safety, Apps, and Social. That means you will have broad exposure to both technical and business problems. You Will: Build and improve the data pipelines, datasets, and internal tools that power decision-making across Roblox Partner with engineers, data scientists, and product teams to turn messy, ambiguous data needs into reliable, reusable data assets Help improve data quality, observability, reliability, and discoverability across the stack Contribute to foundational work that supports multiple product pillars, not just a single surface are

PythonSQLAWSGit
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by providing

PythonJavaGitAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by pro

PythonJavaGitAI
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $137K/yr

Quick readStrong listing-quality and freshness signals

The Code Gen team is tasked with building AI-powered code transformation tools that transform rigid, legacy applications that suffer from poor scalability and high operating costs into modern, microservices-based architectures that are built on top of MongoDB. Join our team and be at the forefront of innovation and creativity. We are looking for a Staff Engineer with domain expertise and years of experience in modernizing legacy applications that are based on traditional database systems. A significant advantage is profound prior experience in leveraging AI, particularly LLMs and GenAI capabilities, to enable reliable, self-driving automation of the code transformation, iterative build, and test processes. In this role, you will be instrumental in initiating technical strategies and ideas, lead the Code Gen team in designing, building, and optimizing our code transformation workflow and tools. You will work on critical components that ensure the scalability, efficiency, and reliability of our services. This involves crafting sophisticated orchestration layers, robust integration points, and high-performance data systems that seamlessly connect and leverage advanced AI capabilities for code generation, build and test. This role will be based remotely in North America. A strong candidate for this position will have Extensive experience (8+ years) in software development and operations, with a proven track record of delivering high performance, correctness, and architectural excellence in fast-paced environments Experience using Relational Databases such as Oracle, MySQL, Microsoft SQL Server or PostgreSQL Experience with tools and methodologies for code analysis, refactoring, and automated testing Experience in designing and implementing complex software systems, collaborating effectively with engineers of all experience levels to achieve high reliability and performance Practical knowledge of integrating GenAI into large-scale, complex systems, including a clear unde

SQLPostgreSQLMySQLMongoDB
🔔

Get new reliability engineer iii jobs in United States by email

Daily job updates · Unsubscribe anytime