Jobs in United States

Lead Staff Software Reliability Engineer Data Platform in United States

2,434 active opportunities · Updated October 2026

Explore current lead staff software reliability engineer data platform jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer, Product Security for our Security Foundations team. Product security sits at the center of the trust customers place in the Snowflake platform, and this role owns the hard technical problems that keep that trust intact. You will drive security architecture and secure-by-design practices across engineering teams, shaping how new products are built rather than reviewing them after the fact. AS A STAFF SOFTWARE ENGINEER, PRODUCT SECURITY AT SNOWFLAKE, YOU WILL: Set the technical direction for product security across multiple engineering teams, influencing architecture decisions before code is written Design and build security-critical services, frameworks, and controls that other engineering teams adopt as defaults Lead threat modeling and security design reviews for the platform's most complex and high-risk systems Partner with product and engineering leaders to embed secure-by-design principles into the development lifecycle Investigate and drive resolution of the highest-severity security issues, then eliminate the underlying class of problem, not just the instance Mentor senior and mid-level engineers, raising the security engineering bar across the organization Own initiatives end to end, from problem framing and scoping through de

PythonJavaAIC++
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the AI Products team The AI Products team is part of the broader Marketplace & Collaboration organization and is focused on bringing AI products on top of Snowflake’s data and application platform to help customers discover, share, monetize, and act on data assets & applications more easily. This team is building the connective tissue of the agentic enterprise: the infrastructure and product surfaces that allow Snowflake customers to seamlessly share datasets, semantic views, and applications, and make them discoverable and executable through Cortex Code, CoWork, and other agentic harnesses. Our strategy is centered on evolving Snowflake Marketplace for the AI era, including packaging data and intelligence into ready-to-use agentic experiences, and enabling governed access patterns that let AI systems safely operate on enterprise data and applications. As a Staff Software Engineer on AI Products, you will Lead the design and delivery of large, complex initiatives spanning multiple teams, turning ambiguous product and platform opportunities into durable technical solutions. Shape the architecture for how datasets, applications, semantic assets, and agentic capabilities are shared, discovered, governed, and invoked across Snowflake surfaces and third-party agent

PythonJavaAIGo
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About Dynamic Tables Dynamic Tables (DTs) are Snowflake's declarative streaming transformation primitive. Customers define a SQL query and a freshness target; Snowflake handles the rest: orchestrating refreshes, maintaining snapshot consistency across a DAG of dependencies, and automatically incrementalizing the computation so that cost scales with what changed. Dynamic Tables is one of the fastest growing products at Snowflake and is a core part of Snowflake’s Data Engineering strategy. The Dynamic Tables performance team is responsible for making incremental refresh fast, predictable, and cost-efficient across increasingly complex query shapes. As a Staff Engineer on this team, you will own the technical direction for critical performance initiatives and be a force multiplier for the engineers around you. What You'll Do Lead the design and implementation of performance improvements to the incremental view maintenance engine, including multi-join incrementalization, novel incrementalization semantics, incremental window functions, and stacked operations. Help define the roadmap for the incremental view maintenance engine, identifying key performance, scalability, and correctness milestones, prioritizing high-impact enhancements, and aligning technical investments with prod

JavaSQLRestAI
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Come join us in this super high impact role as we build the future of Snowflake! Snowsight is the perfect place for data professionals to do their best work, and in this role you'll be making it even better. You'll join a wildly talented, multi-disciplinary team to not only reimagine existing Snowsight experiences, but also to extend it to exciting new surfaces and unlock high value business opportunities. Acting as a startup within Snowflake, together we'll iterate quickly, ship daily, and experiment continuously. This is a phenomenal opportunity to revolutionize how people across the world turn data into knowledge, and knowledge into opportunities. As a senior technical leader on this team, you'll have exceptional impact, both shaping and shipping amazing experiences — and helping the engineers around you do the very best work of their careers. YOU WILL: Act as a technical leader and deep domain expert across the stack, owning an area of the product end-to-end with an eye toward service health, quality, and customer experience. Lead large, cross-team projects from architecture through delivery — identifying opportunities, making the right trade-offs in the context of overall company goals, and driving them to completion. Leverage agentic AI to accelerate your work (and th

JavaReactAIGo
P
📍 United States· Full-time
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Our Merchants organization focuses on onboarding and activation, catalog quality and health, merchant presence (including creator/UGC and trust signals), and ads upsell—helping merchants build healthy catalogs, reach the right audiences, and unlock clear paths to growth. In this role, you’ll lead high-impact engineering at the intersection of commerce platforms, catalog systems, and AI-native experiences. This includes applying ML, AI-assisted workflows, and GenAI where appropriate to improve metadata quality, merchant tooling, and operational efficiency. You’ll partner closely with Engineering Managers, Product Managers, Data Scientists, and other senior engineers to deliver systems with measurable business and user impact. What you’ll do: Own end-to-end technical delivery for cross-team initiatives—from problem framing and technical stra

AWSRestAIGo
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%

From $208.6K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest helps more than 600 million people discover new ideas to help design their life. Our users come to Pinterest to explore ideas and run more than 80 billion search queries every month. Many of these queries represent exploratory search and shopping intent and are broad, which means that the search system should be able to deeply understand this intent, then help people explore content, personalize results, harness visual and multimodal signals effectively, and show the most engaging content up front. All of this means that Pinterest Search presents a unique challenge quite unlike other search systems and the opportunity to innovate on a product that only Pinterest can build. We are looking for an exceptional Senior Staff Engineer to lead Search’s product and front-end strategy across iOS and other platforms such as Android and Web, spann

AWSRestAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $272K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Senior Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Serve as the technical owner for GenAI initiatives within APM, leading design, development, and deployment of ML/AI-powered features across multiple teams. Guide long-term strategy and technical direction for GenAI workflows across APM and related products. Build and benchmark GenAI/ML models using state-of-the-art techniques. Contribute to Datadog’s broader senior engineering community through thought leadership and collaboration on company-wide initiatives. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, sc

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $234K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva

Machine LearningAIGoRust
A
📍 New York, NY, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Who We Are Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise on nearly $9 trillion in assets. Its open platform integrates with nearly 650 software, data and consulting partners to power end-to-end investment operations across firms of all sizes and complexity. Addepar supports clients worldwide with offices in New York City, Salt Lake City, London, Edinburgh, Pune, Dubai, Geneva, Singapore and São Paulo. The Role We are seeking a Staff Full Stack Software Engineer to join the Advisor Experience team as our Technical Lead. Our team is focused on building tools for financial advisors to grow and sustain their business. We oversee bespoke products for advisors and develop advisor-focused capabilities throughout the Addepar platform. In this role, you will be the primary technical anchor for new capabilities including Secure Message Center — a compliant messaging experience built into Addepar's client portal that allows advisors and their clients to communicate directly within the platform. You will partner directly with Engineering Leadership and Product Management to build a modern, scalable architecture from the ground up. Beyond system design, you will act as a true engineering multiplier: setting technical standards, mentoring junior and mid-level engineers, and working alongside other senior engineers and AI specialists to deliver high-impact advisor tools. Applicants must be legally authorized to work in the United States for any employer without requiring current or future visa sponsorship (for example, employment-based visas such as H-1B, F-1/OPT, or similar), and must be authorized to begin work in the U.S. on their first day of employme

JavaReactVueAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

PythonCI/CDLinuxAI
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

C$100K – C$500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to join our CPU design team and lead the front-end RTL physical implementation team. Drive CAD flows on multiple process technologies while working closely with core micro-architects to refine CPU core configurations and optimizing PPA. You’ll work on a CPU based on RISC-V ISA, collaborating with DV, PD, RTL and performance teams to deliver a functional, timing, and power-converged design. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are An expert in physical design practices used to optimize PPA Experienced in high-performance physical design. Proficient in RTL coding (Verilog/VHDL) and familiar with industry-standard tools for simulation and power analysis. Skilled in synthesis, place and route tools including flows and physical design methodology. Background in CPU micro-architecture. What We Need Own front‑end physical implementation and PPA definition for a high‑performance RISC‑V CPU and CPU subsystem Work closely with microarchitects and RTL designers to “make the IP better” by optimizing frequency, power, and area

AWSAISEMHR
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $248K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Data Stewardship team is a group of passionate data practitioners with a diverse background in analytics, data modeling, governance, compliance, and scaled data quality. We are responsible for ensuring Airbnb is meeting its compliance obligations across our data ecosystem and ensuring data consumers are able to easily identify the best data for their needs. We support the pipelines, programs, and policy-bodies that make this possible. You’ll be part of the overall Data Infrastructure organization that is responsible for online and offline data infrastructure across the company, and the components that transition data between these environments. The Difference You Will Make Set the North Star: you will define the multi-year vision for Data Governance and Data Quality that scales with our global business and evolving AI landscape. Organizational Influence: Act as a primary consultant for executive leadership on data governance, ensuring that compliance and stewardship are integrated into the data product lifecycle from day one. A Typical Day Architect Ecosystems: Lead the design of overarching data architectures that don't just solve today’s batch needs but anticipate future real-time and AI-driven requirements. Scale Best Practices: Rather than just "ensuring quality," you will define the best practices, tooling, and culture that enable data organizations to maintain high-quality data autonomously. Policy Stewardship: Lead cross-functional task forces (Infosec, Legal, Privacy) to navigate complex regulatory landscapes (like GDPR or AI Act) and translate them int

JavaSQLAIGo
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $244K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: AI and ML are at the heart of the Airbnb product. From Trust to Payments, and from Customer Service to Marketing, we rely on ML to ensure that guests and hosts have the best possible experience with Airbnb. The Core ML team is responsible for driving CSxAI (Customer Support x Artificial Intelligence) initiatives by adopting Generative AI technologies to enable an intelligent, scalable, and exceptional service experience. The team develops and enhances AI models, ML services, and tools including LLM fine-tuning and optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning, and guardrails for a wide range of applications at Airbnb. The richness of Airbnb's data, the complexity of its marketplace, and the variety innate in our product mean that we need to operate at the state of the art of AI practice. We are committed to long-term innovation to solve complex problems, and to do that we need experienced ML ​​The Difference You Will Make: In this Senior Staff role, you will set technical direction and lead execution for ML evaluation and the end-to-end data flywheel powering CSxAI products (e.g., assistive agents, issue resolution, and tooling). Your work will define how we measure quality, how we turn feedback into learning signals, and how we continuously improve models and products safely and efficiently. You will partner closely with product, engineering, design, operations to build evaluation systems that are trusted, scalable, and actionable - connecting offline metrics to online outcomes. A Typical Day: Define evaluation strategy and suc

AgileAIGoRust
S
📍 Portage, Michigan, United States
✓ High-confidence listingCompany trend +364.7%
Quick readStrong listing-quality and freshness signals

Work Flexibility: Onsite It's Time to Join Stryker! Stryker is seeking a Staff Systems Engineer, Robotics to lead systems engineering activities for critical electromechanical components within a complex robotic medical platform. In this role, you will serve as a systems engineering owner for robotic end effectors , including powered surgical tools and future attachments that interface directly with the robotic system. You will help translate user and clinical needs into system architecture and requirements, define interfaces across electrical, mechanical, and software disciplines, and guide the product from requirements development through integration and verification. This role requires a highly collaborative engineer who can bring together input from multiple technical disciplines and maintain a system-level view throughout development. What You Will Do Systems Engineering Own systems engineering activities for robotic end effectors and related electromechanical subsystems. Translate user, clinical, and product needs into clear, verifiable system requirements and design inputs. Develop and refine system architectures, interfaces, and functional requirements across electrical, mechanical, software, and controls disciplines. Allocate and decompose system requirements to appropriate subsystems and engineering disciplines. Lead technical trade studies, design assessments, performance analyses, and other systems engineering activities used to guide design decisions. Support concept development and architecture definition for new end effectors, features, and product capabilities. Requirements, Integration, and Verification Establish and maintain requirements traceability throughout the product development lifec

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, cloud services, mechanical engineering, electrical engineering, and product design to deliver reliable, production-ready devices at scale. Within Consumer Devices, Hardware Engineering eXperience, or HEX, is a new bootstrapped team building the environments, applications, compute, product-data systems, and workflows that let hardware engineers do their work without needing to troubleshoot the machinery underneath. HEX owns virtual engineering environments, HPC/GPU compute, storage, networking, licensing, MCAD/ECAD/CAE applications, PLM, product data, automation, validation, and support as one connected system. About the Role As a Staff PLM & Engineering Applications Engineer, you will be one of the first technical builders of HEX and the primary counterpart to the HEX lead. You will own the engineering-application and product-data side of the hardware engineering experience, with an initial focus on NX, Teamcenter, licensing, parts import, integrations, packaging, validation, and user workflows. This is not a traditional Teamcenter administration role and not a Corporate IT application-support role. You will take complex, fragile workflows and turn them into reliable engineering systems. This role is highly hands-on and systems-oriented. You will not inherit a mature environment and support queue. You will help build a fresh one, replacing manual setup guides, tribal knowledge, repeated support issues, and team handoffs with tested automation and reliable workflows. In This Role, You Will Own the technical architecture, deployment, configuration, integration, validation, and long-term operation of NX and Teamcenter. Build reliable workflows for parts import, product-data migration, metadata quality, BOMs, revisions, lifecycle states, and releas

PythonAWSRestAgile
🔔

Get new lead staff software reliability engineer data platform jobs in United States by email

Daily job updates · Unsubscribe anytime