Jobs in United States

Cost Operations Lead in New York

27 active opportunities · Updated October 2026

Explore current cost operations lead jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $204K/yr

Quick readStrong listing-quality and freshness signals

The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort

KubernetesGitAIGo
W
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $125K/yr

Quick readStrong listing-quality and freshness signals

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: WPP’s Data & Technology Solutions is the WPP’s unified global data products and technology team. We work with our agencies and clients to build data-driven solutions and technology product to power marketing transformation. WPP Open is our AI platform for marketing, it connects marketing professionals, data, tools and AI in a single place. WPP Open is the simplest, safest and fastest way to realize the benefits of scaled AI – delivering better-informed creative ideas faster, at scale and at lower cost. We’re endlessly curious and our team of thinkers, builders, creators and problem solvers are over 2,000 strong, across 20 markets around the world. We are seeking a strategic and commercially driven leader to represent and execute WPP’s data partnership strategy with Media Owners and Platforms. This role will be critical in scaling local markets Data Strategies Globally You are: A data leader with strong experience in partnerships and platform ecosystems with a robust knowledge of the AdTech industry Commercially minded, with experience negotiating and managing strategic par

AIGoRustMarketing
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio

PythonSQLAWSCI/CD
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. Today, CLEAR is well-known as a leader in digital and biometric identification, reducing friction for our members wherever an ID check is needed. We’re looking for a Senior Software Engineer to establish our Observability framework and foundations. You will join us to accelerate building and scaling our innovative systems that support our growing identity platform. You will drive on Observability best practices to find and fix gaps in our observability and our overall systems. You will also lead practices such as load testing, capacity planning, game days, chaos testing, and incident post-mortems. What You Will Do: Embed within the Engineering pillar to deeply understand the product and implement observability across all key flows Facilitate and build load testing cases, ensuring we understand the limits and scaling factors of our services and systems Contribute to observability and support the design of new services and systems, ensuring highly reliable and scalable concepts are implemented Build and lead practices such as game days, chaos engineering, and failure analysis Build long-term capacity plans, with an eye toward reliability and cost-efficiency Who You Are: 6+ experience writing production-grade software in a modern language, such as Java and Python. Strong knowledge of distributed systems concepts (think CAP theorem), microservices architecture, and distributed tracing . Experience with modern observability systems such as Datadog. Experience with performance debugging tools and patterns. You should be able to read a f

PythonJavaGitRest
P
📍 New York, NY, United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Storage Services team operates Pinterest's scalable online structured data storage platform—supporting both SQL-based table models and complex graph data structures—managing 100TB+ datasets and serving over 1.5M queries per second across many of Pinterest's most important products. We're looking for an exceptional Staff Software Engineer to lead the technical strategy and execution of our storage infrastructure initiatives, defining how these systems are designed, built, and operated. You'll drive innovation across distributed SQL, high-throughput/low-latency query processing, graph workloads, and the developer experience for storage clients. What you’ll do: Provide technical guidance and direction to a high-performing team building reliable, performant, and cost-efficient storage systems that operate at massive scale and power busines

PythonJavaSQLAWS
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Most of the value of owning a model shows up at serving time. We're building a platform that covers the whole life of an LLM -- train it, deploy it, observe it -- and inference is where teams feel the difference every day. We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell. You will do hands-on inference research at Modal, working with the research lead to pick high-impact bets and owning them end to end. The bets that matter most are the ones that move cost per token and tail latency on the workloads our customers actually run. What you'll do: Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spik

P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The FinOps function is responsible for financial accountability, visibility, and optimization across all engineering-related spend at Plaid. This includes cloud infrastructure, AI/ML and data workloads, third-party SaaS tools, and other technical investments that support Plaid’s products and internal platforms. The team operates at the intersection of Engineering, Product, and Finance, ensuring that spending decisions are transparent, intentional, and aligned with product strategy and business priorities. Rather than functioning as a cost-control or approval layer, FinOps enables teams to understand, own, and optimize their spend while maintaining engineering velocity. Responsibilities Monitors and analyzes engineering spend across cloud, AI/ML, data platforms, and SaaS, identifying trends, anomalies, and optimization opportunities. Builds and maintains forecasts for engineering spend, partnering with Finance and engineering leaders to understand drivers, assumptions, and risks. Partners with engineering, product, and TPMs to incorporate cost considerations into roadmaps, architectural decisions, and execution plans. Leads cost optimization initiatives, such as rightsizing, commitment strategies, an

SQLAWSAzureGCP
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $244K/yr

Quick readStrong listing-quality and freshness signals

We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—allowing for seamless collaboration and problem-solving among Dev, Ops and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Team: As organizations rapidly invest in AI applications and build out AI labs, telemetry volumes are growing exponentially and costs are becoming unpredictable. From LLM interactions to agentic workflows, AI systems generate unpredictable streams of logs, driving up costs and making it harder to maintain efficient observability. These challenges are critical for organizations in regulated industries with strict data residency requirements, where data must remain within controlled environments. Datadog’s Bring Your Own Cloud (BYOC) team is reimagining what observability and security look like at petabyte scale in the AI era. The Opportunity: The Group Product Manager - Bring Your Own Cloud (BYOC) role is responsible for defining and bringing to market the next generation of telemetry analytics and insights capabilities in an AI-first environment. This role is highly technical and creative in nature as you will envision novel ways to enable customers to cost-effectively explore, analyze and report over petabytes of data through a welcoming and easy-to-use interface. You will partner with various teams to take advantage of BitsAI capabilities and surface critical insights on volume usage and retention for popular use cases. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead and grow a team

SQLAWSAzureRest
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $276K/yr

Quick readStrong listing-quality and freshness signals

Team description At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the platform that connects these agents to Datadog: the MCP Server, the tools and retrieval surfaces agents call into, and — critically — the evaluation systems that tell us whether an agent's experience on Datadog data is actually getting better over time. This role is about that last piece. We're hiring a Staff Applied Scientist to define what "good" means for an Agentic interface at Datadog and to build the measurement systems that make it true. "Good" isn't one number — it spans answer quality, tool-selection accuracy, retrieval relevance, latency, token cost, and end-to-end agent success on real customer workflows. You'll design the evals, build the datasets, define the metrics, and partner with the AI engineers on the team to land the platform that lets every product group at Datadog ship integrations that are demonstrably better release over release. The space is full of open research questions. How do you evaluate an agent end-to-end when the trajectory is non-deterministic? How do you score tool selection when the tool catalog has hundreds of entries and grows weekly? How do you build a measurement system that catches regressions across first-party and third-party agents at once, without each team writing their own harness? If those are the problems you want to spend your time on, come build this with us. Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply. What You’ll Do: Own the evaluation strategy for Datadog's AI agent integrations. Define the metrics — offline and online, quali

AIGoRustSpring
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%
Quick readStrong listing-quality and freshness signals

At Datadog, we’re on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Observability Data Platform (ODP) is the backbone of everything Datadog delivers – powering how data is ingested, stored, routed, and surfaced across every product at planet scale. As a Senior Product Manager for ODP, you will work with world-class engineers and cross-functional partners to shape how the platform is deployed, controlled, and operated. You will define product direction across the control plane and data layer, translate complex infrastructure trade-offs into clear roadmap decisions, and help customers get the most from their observability investment – regardless of architecture, topology, or scale. At Datadog, we place value in our office culture – the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You Will Do: Develop a deep understanding of the Observability Data Platform customers – platform engineers, SREs, and product managers that own the product verticals – their infrastructure challenges, deployment topologies, and cost-to-serve trade-offs. Define product direction across multiple ODP surfaces, including the control plane and data layer, by articulating clear problem statements and desired outcomes, and partnering with engineering on technical approach and sequencing Lead conversations with design partners and strategic customers to understand real-world platform pain points, validate product assumptions, and guide solutions from early prototypes through General Availability Develop a co

AIGoRustSpring
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $89K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for someone to join the Datadog Procurement team and help expand the Strategic Sourcing group. Make an impact by being a trusted analyst in several buying categories and continue to prove the value that Strategic Sourcing brings to the organization. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Leverage AI tools (e.g., Claude, ChatGPT) to enhance market research, accelerate data analysis, and improve efficiency in sourcing workflows Identify opportunities to incorporate automation and AI into procurement processes to drive scalability and smarter decision-making Support and execute sourcing strategies directed by sourcing managers with a specific focus in G&A categories (People Team, Finance, Legal, Real Estate) Develop strong relationships with stakeholders and assist teams to independently evaluate their best practices and adopt a value-driven mindset when purchasing on behalf of the organization Lead and execute Sourcing events (RFx) Negotiate pricing and other business terms with vendors on net-new purchases and renewals Manage the renewals list for designated categories and ensure accuracy of information and proactive communication Analyze spend data and identify potential cost savings opportunities Create pricing models based on vendor proposals to quantify various buying scenarios Build license forecast models and utilization analyses for enterprise software renewals to help uncover usage and savings opportunities Collaboratively align with adjacent functions like FP&A on utilization and forecast models Operate effectively in ambiguous or evolving problem spaces, helping bring structure, clarity, and forward momentum to loosely defined sourcing initiatives Monitor market trends that impact spend categories and

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $320K/yr

Quick readStrong listing-quality and freshness signals

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

GitMachine LearningAIGo
🔔

Get new cost operations lead jobs in New York, United States by email

Daily job updates · Unsubscribe anytime