Jobs in United States

Staff Software Reliability Engineer Data Platform in New York

81 active opportunities · Updated October 2026

Explore current staff software reliability engineer data platform jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit

GitRestMachine LearningAI
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for strong backend engineers who love building a developer tools used by the largest AI companies in the world. You’ll be building for things at scale, but also for new AI workflows that change every day. Requirements: Experience building and shipping modern web applications end-to-end. We care more about what you’ve built than how many years you’ve been building. Comfort working across the stack: TypeScript on the frontend, Python services on the backend, and ClickHouse for data and analytics. Deep knowledge of observability tools and patterns used for large-scale workloads such as custom sandboxes, training and inference for large language (LLM) and diffusion models. Experience with at least one of: billing/payments systems, B2B SaaS tooling, or enterprise software, or LLM / diffusion models inference and training loads. Strong product instincts; yo

TypeScriptPythonAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

This Engineering Manager will lead the Data Visualizations Explorations team within the Graphing organization, setting product direction, and coaching and developing team members. They will staff and drive projects that build end to end experiences for Datadog’s core users: observability engineers. This includes extending the capabilities of core widgets like Hostmap and Geomap Visualizations and finding innovative ways to leverage existing Datadog data sources. This role also involves close partnerships with other product teams to deeply understand customer needs and deliver compelling data experiences. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Work with Product Management and Design to plan and staff projects for the Data Visualization Explorations team Coach and develop engineers at various levels Ensure strong cross-team communication and design best practices around project management, code review, architecture patterns, and more. Proactively anticipate cross-team dependencies and blockers to goals. Identify opportunities to appropriately reuse or customize features across dashboards, notebooks, and product pages. Deeply understand the needs of our customers and other Datadog products we work with. Participate in customer conversations, review product briefs, read feature requests, and coach team members to adopt these practices as well. Define and maintain high standards for operations practices, including bug triage and remediation, incident response, and gathering and analyzing performance telemetry for our widgets. Who You Are: At least 2 years of people management experience in a software engineering or similar setting Strong TypeScript/JavaScript skills, including familiarity with front

JavaScriptTypeScriptJavaReact
DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

PythonAWSAzureCI/CD
N
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -40%
Quick readStrong listing-quality and freshness signals

About Novo Small businesses are the backbone of the US economy — nearly half of GDP and the private workforce — yet big banks don't give owners the access, assistance, or modern tools they need to grow. We started Novo to change that. Our mission is to increase the GDP of the modern entrepreneur by building the go-to banking platform for small businesses. Novo is a fintech, not a bank. We give entrepreneurs, freelancers, and SMBs an operating system for their money: no-fee business checking, debit and credit cards, loans and funding, payments across every rail (ACH, wires, FedNow), invoicing, bill pay, expense management, and integrations with the tools they already run their business on. We've helped hundreds of thousands of businesses get powerfully simple banking, and we're just getting started. Why Novo? A mission-driven team building for the people big banks pretend don't exist Real ownership: we make decisions that affect real businesses every day, and we trust people to act on that A flat culture where ideas win on merit — hierarchy is not authority Modern stack, modern tooling, and a team that takes AI seriously as part of how we work Competitive salary and stock options for every full-time employee Medical, dental, and vision coverage Learning and development budget Offices in NYC and India, with Wednesday team lunches that rotate cuisines The Role We're looking for a Staff Product Designer to be the highest-leverage individual contributor on our product design team. You'll own the hardest, most ambiguous design problems at Novo — the ones that cut across checking, payments, credit, and invoicing, and across web, iOS, and Android — and you'll set the bar for what great looks like for everyone else. This is not a management role. It is a role with real authority over the quality and coherence of the product. You'll partner directly with product leads, engineering leads, data, and compliance to shape what we build, not just how it looks. You'll multiply the t

ReactAISwiftGo
T
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -50%

From $170K/yr

Quick readStrong listing-quality and freshness signals

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday. About the Role Machine Learning is a cornerstone at Taskrabbit, and we're looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale. Taskrabbit's greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There's significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expan

PythonSQLDockerKubernetes
T
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -50%

From $170K/yr

Quick readStrong listing-quality and freshness signals

About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role is hybrid requiring 2 days in office at our San Francisco or NYC hub every Tuesday & Wednesday. About the Role Machine Learning is a cornerstone at Taskrabbit, and we're looking for a Staff Machine Learning Engineer to join our team and lead the next phase of our customer retention strategy. This is a critical, full-stack role for an individual who is passionate about the end-to-end lifecycle: from initial research and model development to building the robust systems that power repeat customer engagement and lifetime value growth at scale. Taskrabbit's greatest growth opportunity lies in deepening customer relationships and accelerating repeat purchases. Our most valuable customers are those who return frequently, discover new service categories, and increase their spending over time. There's significant untapped potential in the marketplace: repeat customers spend 3-5x more than one-time users, and category expan

PythonSQLDockerKubernetes
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. Our Credit team at Plaid is building the largest cash flow based Consumer Reporting Agency (CRA) in the US to deliver lending solutions across the full lender lifecycle — including underwriting, verification, and servicing. We help lenders make faster, more informed decisions using consumer-permissioned financial data. As the Product Manager for our CRA Network team, you will grow the number of consumers that have permissioned their data to our CRA, Plaid Check. You will build a more robust experience for consumers to interact with the CRA and learn more about sharing cash flow data with lenders. Responsibilities Own and accelerate the growth and health of our CRA network Drive credit initiatives across the broader Plaid network Own consumer consent UX experience including conversion and improvements Build trust through our CRA consumer experience Qualifications 5+ years of product management experience Must have built and owned a lending product at a FinTech Experience working on a growth product - especially on a network product Deep understanding of the E2E lending process from acquisition to originations to servicing Excellent communication skills and the ability to advocate f

AWSAIRustExcel
P
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest is uniquely positioned across the full consumer journey—from inspiration and discovery through consideration, planning, and action. Creators bring ideas to life, Pinners signal emerging interests and intent, and advertisers help people act on what inspires them. This role will define how Pinterest measures that flywheel, helping advertisers understand how Pinterest and its creators build awareness and drive business outcomes, while helping creators understand the value they generate and unlock new opportunities with brands. The ideal candidate is a Staff-level product leader with substantial ads measurement experience and a deep understanding of how advertisers set objectives, deploy tactics across the full funnel, and evaluate outcomes. They have worked cross-functionally with engineering, data science, research, PMM, sales, an

AWSRestAIGo
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team. As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity. What excites you Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling. Architect and manage the SLO and error-budget

AWSKubernetesAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $232K/yr

Quick readStrong listing-quality and freshness signals

Datadog is entering a new chapter in how our product looks, feels and operates. We’re building a small, high-leverage Design Lab team to define that evolution, and we’re looking for a Staff Visual Designer to help author it. Our design system powers everything we ship. Now we’re ready to evolve it. As Datadog expands into AI-driven experiences and more complex product surfaces, the visual language needs to grow with it. We’re looking for someone who can help define that next phase, someone with taste, judgment, and the confidence to shape how the product feels at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do Define and evolve Datadog’s visual language across product surfaces. Establish the visual quality bar for the platform. Lead exploratory and concept work that pushes our system forward. Translate bold visual ideas into scalable patterns that can live inside a design system. Partner deeply with Brand, Product Design, and Design Systems to unify expression across product and marketing. Influence executive stakeholders through clear, compelling visual storytelling. Prototype, test, and refine new patterns before they become systemized. Help shape how visual direction is integrated into our design process. You will report directly to the Design Director and work as part of a small, focused team defining the future state before it scales across hundreds of designers and engineers. Who You Are You have 10+ years of experience in visual and/or product design You have led a meaningful visual evolution or rebrand within a mature product ecosystem. You have strong taste and are comfortable defending a point of view. You know how to translate brand into scalable product systems. You think in systems, but you’re motivated by craft. Yo

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $276K/yr

Quick readStrong listing-quality and freshness signals

The Dashboards product is Datadog's unified single-pane-of-glass for metrics, logs, and traces—a comprehensive treasure trove of observability data. We are transforming Dashboards into an AI-native control surface and the central hub where every team moves seamlessly from question to insight to action – providing a guided experience that feels like having an expert SRE at your side and ensuring the entry point is never an empty canvas. We're hiring a Staff Applied Scientist to define and guarantee the quality of this AI system at scale. "Good" isn't one number — it spans answer quality, tool-selection accuracy (critical given the growing catalog of data sources and visualizations), retrieval relevance, latency, token cost, and end-to-end agent success. The space is full of open questions. How do you evaluate an agent end-to-end when the trajectory is non-deterministic? How do you score tool selection when a user’s query can result in the agent making decisions against dozens of visualizations and data sources – both of which are growing month over month? How do you build a measurement system that catches regressions across all widget types and data sources (e.g., enforcing correct grouping, sorting, and time overrides), and is easy to use and extend by dozens of teams? If those are the problems you want to spend your time on, come build this with us. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the evaluation strategy for Dashboards, as well as sister teams within our organization. Define the metrics — offline and online, quality and cost, single-turn and multi-turn — that the team and the broader organization optimize against. Build the eval datasets, golden traces, and regression harnesses that catch quality changes before they hit customers, an

AIGoRustSpring
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $204K/yr

Quick readStrong listing-quality and freshness signals

The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort

KubernetesGitAIGo
🔔

Get new staff software reliability engineer data platform jobs in New York, United States by email

Daily job updates · Unsubscribe anytime