Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for Members of Technical Staff to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. You may be a good fit if you have: 5+ years of engineering experience running production infrastructure at a large scale Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters Experience with Kubernetes dev and production coding and support Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving Experienc
Jobs in United States
Staff Software Reliability Engineer Data Platform in United States
1,760 active opportunities · Updated October 2026
Showing
15 jobs
Explore current staff software reliability engineer data platform jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Issue Workflow is Sentry's primary product surface. Our issue platform processes billions events daily and turns them into actionable insights that help millions of developers fix bugs faster. As a Staff Software Engineer on the Issue Workflow team, you'll architect the systems that power this experience. You'll work at the intersection of high-scale distributed systems and product engineering, building real-time data pipelines, search backends, and analysis systems that surface signal from noise. This is product engineering at massive scale—where every architectural decision impacts millions of debugging sessions. You'll be the technical leader who shapes how Sentry groups issues, how we make search lightning-fast, how we enable sophisticated agentic workflows, and how we ensure that the product is performant even at billions-of-events scale. Your work will define what's possible for the most trafficked part of Sentry's platform. In this role you will Drive technical strategy and roadmap. Partner with engineering leadership, product, and design to shape the multi-quarter technical vision for Issue Workflow platform. Make strategic calls about architectural direction, technology choices, and technical debt. Ensure the team is building a strong foundation to scale with Sentry's growth. Solve complex performance and scalability challenges. Champion product quality and user experience. Build features that don't just work—they delight. You understand that milliseconds matter in the developer experience. You sweat the details of interfaces, error messages, loading states, and edge cases. You instrument everything s
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role This isn’t a typical engineering role. You won’t be embedded in a single product team or siloed in one product area. Instead, you’ll sit within Platform Engineering, own the AI-assisted coding domain, and work across all of engineering at Sentry, focused specifically on how AI coding agents participate in our software development lifecycle. For AI coding agents to work well in our repo, the internal systems they depend on need to be accessible via API, not locked behind UIs that require human interaction. Right now, many of those systems aren’t agent-ready. You’ll audit and prioritize that gap, expose those systems programmatically, and build the connections that let tools like Claude Code operate on them end-to-end. From there, the scope expands to improving the quality of AI-generated pull requests and automating the engineering work that’s important but consistently deprioritized. You will look from context engineering standpoint to see what to send to our model; you will look from harness engineering standpoint to see the tools it can use, the permissions it has, the state it carries forward, the tests it has to pass, the logs you capture, the retries, checkpoints, guardrails, and evals. You’ll work closely with the dev infrastructure team as your home base, then collaborate across every product team coding in our repo once the tooling foundation is in place. It’s a broad role with real impact, and the work you do will directly change how Sentry engineers ship software. What You’ll Do Audit Sentry’s internal developer systems and make them API-ready for AI agents. You’ll prioritize and drive the work of ex
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role This is a zero-to-one Product Engineering role on Replit's Money team. You'll bring our financial partnerships alive from scoping to integrating and operating them so that the builders on Replit, and the Apps and Agents they ship can transact reliably anywhere in the world. Money at Replit isn’t traditional payments. Builders are publishing apps that earn revenue. Agents are spending and being paid for work. New protocols for agentic commerce now exist to power the future of commerce: Shared Payment Tokens, the Agentic Commerce Protocol (ACP), the Universal Commerce Protocol (UCP), the Machine Payments Protocol (MPP). Replit is one of the platforms that will define what they look like in practice and lower the barrier to entry. To make any of that real, we need someone who can sit between Replit engineering and our financial partners including billing platforms, payment processors, agentic-commerce protocol partners, tax and compliance vendors to turn signed contracts into live, reliable integrations. You'll be the engineer partners ask for, and the engineer the rest of Replit relies on when a new monetization surface needs to ship. This role is a fit if you like writing code and you also like being in the room when a partnership is being scoped, because you know that the design decisions made in that room are the ones that bind the integration for years. You will Take financial partnerships from zero-to-one: scope the integration with the partner, pressure-test data and protocol specs, design the system, build it, ship it, and operate it. Own the technical relationship with Replit's financial partners across the full lifecycle: billing platforms, payment processors, agentic-commerce protocol partners, t
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: This is a Product Engineering role specialized in monetization systems, and will lead that space at Replit. It’s a direct line to business impact. But it’s also critical to get right for our users. These are some of the most critical user journeys to get right. Getting them wrong creates the most frustrating experiences for users. So we’re looking for engineers who can build reliable and scalable billing systems and abstractions, while also translating that to an intuitive and friendly user experience. You will: Lead the architecture and implementation of monetization systems at Replit. Create seamless payment experiences for users for both product-led and sales-led motions. Build new abstractions and APIs for other engineers at Replit to monetize their new products. Iterate on pricing and packaging tactics to drive revenue growth. Examples include coupon codes and referral systems. Create monitoring and feedback systems so that we can proactively spot problems, fix them, and optimize performance. Required skills and experience: 6+ years of engineering experience, with strong skills working on the backend. Direct working experience in at least one of the following: Subscription platforms Usage-based billing SaaS Taxation Payment platforms Tokenization Self-directed and comfortable working autonomously in ambiguous environments. Excellent problem-solving skills with ability to debug complex billing issues and edge cases. Experience implementing customer-facing billing interfaces that simplify complex pricing structures. Tools + Tech Stack for this role: Python, TypeScript, React, Postgres, GraphQL, and Nodejs. Bonus Points : Experience working with Orb, Metronome, or Stripe usage based billing. Experienc
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. Replit is building the world’s most accessible AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it’s an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit builds complete apps better than anybody thanks to our full suite of services that handle app integrations, storage, hosting, analytics, and more. We don’t just build apps in development, we handle the full lifecycle into production and beyond. About the role: Help power the development of Replit Agent as a technical leader for the Replit Cloud organization. You will report to the Vice President of Engineering. The Replit Cloud team builds Replit’s first party cloud infrastructure so users can build, scale, and succeed entirely on Replit. They manage databases, application storage, app publishing and hosting, development/production environment splitting, custom domains, and more. By having a set of first party services that integrate seamlessly, you will power one of Replit’s key product differentiators. You will: Help lead major projects, either by taking new products from 0->1 or doubling down on our first party primitives to keep winning users. Work closely with designers and product managers, to quickly iterate on Replit Cloud to continually grow and improve the product. Identify the hardest technical and/or quality problems holding us back, and then build solutions. Mentor and develop new senior engineers to help grow the team. Ship product and build infrastructure as a true full stack builder using: TypeScript, React, CSS, Postgres, Go, and Terraform. Examples of what you could do: Leverage our unique cloud infrastructure to build diffe
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About Replit Replit is building the world’s most ubiquitous AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it’s an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit is also the leader in secure vibe coding. We protect apps, give users features to manage security risks, and help them vibe code more safely. About the Role In this role you will build powerful tools that help product engineers iterate rapidly on the Agent experience and directly enhance the core Agent itself. You’ll bridge the gap between the AI team (working on the core Agent logic) and the UX team (crafting delightful Agent experiences), enabling both groups to excel within their specialties. This role blends systems engineering, developer experience and product engineering. We tackle complex challenges across the full stack, from browser-based interfaces to high-performance backends to Linux systems engineering. We’re looking for engineers who have a keen sense of the product experience and how to power it with performant systems. On this team, you’ll have the opportunity to grow your skills across our infrastructure and product, and to lead end-to-end efforts with meaningful impact. We value diverse perspectives and encourage candidates from all backgrounds and experiences to apply. You Will Build high-throughput backend applications and services, like streaming chat between user and agent. Design a collaborative "Multiplayer Computer" that lets humans and AI agents work together on shared shells, filesystems, and state—conflict-free and in real time. Develop infrastructure (frontend & backend) that empowers product enginee
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: As a Staff Product Engineer at Replit, you’ll work closely with other product and platform engineers, designers, sales representative, and product managers to build features that help users collaborate with their team to go from idea to software fast. You’ll be at the forefront of shaping and experimenting on what our tens of millions of users love. You will: Help lead major projects and take new products from 0->1 Identify the hardest technical and/or quality problems holding us back, and then build solutions Chart high level technical direction and follow up to make sure those projects come together to deliver on results Mentor and develop new senior engineers to help grow the team Ship new features and build infrastructure using: TypeScript, React, CSS, GraphQL, Node.js, and Postgres Required skills and experience: A minimum of 7 years of professional software development experience Experience in a technical leadership role, working cross functionally Working experience building full stack applications with TypeScript Working experience building directly for users Bonus Points : You’re excited about the future of programming and have experience working with IDEs, terminals, or other common developer tools You’ve had previous experience working at a startup in a cross-functional engineering role This is a full-time role that can be held from our Foster City, CA office. The hybrid role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Making data driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide tooling and guidance to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. Engineers on Data Infrastructure are domain experts in Data Warehouse, Data Lakehouse, Spark, Workflow Orchestration, and Streaming technologies. We scale our existing data pipelines in a performant and cost efficient way while creating the necessary abstractions to make developing on top of this platform extremely simple for other engineers at Plaid. Responsibilities Contribute towards the long-term technical roadmap for data-driven and machine learning iteration at Plaid Leading key data infrastructure projects such as improving ML development golden paths, implementing offline streaming solutions for data freshness, building net new ETL pipeline infrastructure, and evolving data warehouse or data lakehouse capabilities. Working with stakeholders in other teams and functions to define technical roadmaps for key backe
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. AI and intelligent systems are driving the fifth paradigm shift, following previous technological revolutions like mainframes, personal computers, the internet, and mobile devices. We believe, in the foreseeable future, AI will revolutionize the FinTech industry - from how consumers understand and manage their finances, to how developers build applications and how all companies operate. The fintech industry landscape will undergo a fundamental reshape. Plaid in the FinTech AI Ecosystem Plaid is uniquely positioned to become the financial data and insights backbone for AI applications and platforms in this evolving ecosystem. We believe consumers should be able to understand and manage their financial life through conversational AI interfaces using natural language. We believe consumers should have peace of mind with a trustworthy consent and authorization manager when agents shop for them. We believe identity verification and financial fraud prevention in AI-powered products should feel seamless and embedded for the end users. The list goes on. The most important AI companies, major fintechs, and customer agent platforms are actively trying to integrate Plaid into AI-powered products and solutions t
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The Data Governance team makes sure Plaid handles consumer and customer data responsibly — and can prove it. Our mission is to enforce Plaid's privacy commitments and regulatory obligations in the systems themselves rather than in policy documents: we build the platform and controls that govern how data flows through Plaid — where it lives, who can use it, for what purpose, and for how long. That includes verifiable deletion of consumer data on request, enforcement of data-use restrictions so downstream systems can only use data in permitted ways, and the cataloging and classification that let Plaid know what data it holds and how sensitive it is. We operate at the scale of Plaid's entire data footprint, and correctness and auditability matter to us as much as throughput. As a Staff Software Engineer on Data Governance, you will set the technical direction for how Plaid enforces data governance at scale. You'll lead the design of distributed backend systems that reliably delete, restrict, and track data across dozens of services, making architectural decisions whose blast radius spans the whole company. You'll drive multi-quarter initiatives from ambiguous privacy and regulatory requirements through
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Security Engineering is the engineering function inside the Plaid security org that focuses on developing the industry-leading security systems and infrastructure. Security Engineering owns most of Plaid’s security-related infrastructure: secure data storage, key management systems, internal identity platform, internal authentication systems, internal permission management, and internal authorization service. We develop solutions across data encryption, key management, access control, and data loss prevention to protect sensitive consumer data. We believe in the Zero Trust security model and are always looking for ways to improve our authentication and access control platforms. About the role: You will develop security capabilities to secure Plaid infrastructure and sensitive data access. You will lead the team’s strategic planning in collaboration with the manager and other senior engineers. You will own, maintain, and build Plaid’s security infrastructure and services like IAM Gateway, Key Management System and Network Firewall. You will consult with product engineers to ensure Plaid services meet security standards. You will help educate and support other engineering teams to improve security in
From $196.9K/yr
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. The opportunity The Autonomous Freight Systems team is a brand new, AI-first engineering team in San Francisco. This team will own Flexport’s client-facing rates platform and the self-serve freight booking experience—two of the highest-leverage surfaces in our Client App that dictate how clients see pricing and book freight without manual intervention. As a Staff Engineer, you will be the technical anchor for our next-generation AI-powered rates platform. We aren't just building a UI; we are building AI agents that handle real logistics work: parsing complex rate sheets, managing pricing intelligence across ocean, air, and trucking, and making "Self-Serve" a reality for thousands of shippers. You will build the intelligence layer that allows clients to commit freight on technology alone, with no account executive and no operations touch. This is a ground-floor opportunity to shape technical direction, set up a new codebase, and build the platform that moves Flexport from an assisted-sales model to a tech-run one for the long tail of our client base. You will partner with Pricing, Sales, Ops, and Design on hard domain problems and ship to the most-used client
From $182K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team GoDaddy is hiring a Staff Software Engineer to help define and scale our Internal Developer Platform —a centralized system that powers how engineers across the company build, deploy, and operate software. This platform is used by 1,000+ engineers to manage everything from cloud infrastructure and application lifecycle to security, compliance, and cost transparency. In this role, you’ll operate as a technical leader at platform scale, shaping the architecture and direction of systems that directly impact engineering velocity across the company. You’ll work on high-impact initiatives such as AI-powered developer tooling, next-generation API platforms, and the evolution of a unified developer experience spanning APIs, CLI, and UI. This is a high-ownership, high-visibility role where Staff Engineers drive decisions, influence product direction, and partner across infrastructure, security, and platform teams. If you’re motivated by building systems that improve how other engineers work—and want to have a measurable impact on developer productivity at scale—this team sits at the center of GoDaddy’s engineering ecosystem. What You’ll Get to Do... Design and build platform services that power GoDaddy’s internal developer ecosystem, used by 1,000+ engineers. Lead architecture and technical direction for high-scale APIs, infrastructure orchestration,
Other cities to consider
More places hiring for this role
Get new staff software reliability engineer data platform jobs in United States by email
Daily job updates · Unsubscribe anytime