Jobs in United States

Design And Cost Estimation Head in New York

234 active opportunities · Updated October 2026

Explore current design and cost estimation head jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $204K/yr

Quick readStrong listing-quality and freshness signals

The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort

KubernetesGitAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $244K/yr

Quick readStrong listing-quality and freshness signals

Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,

AWSAzureGCPKubernetes
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%
Quick readStrong listing-quality and freshness signals

At Datadog, we’re on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Observability Data Platform (ODP) is the backbone of everything Datadog delivers – powering how data is ingested, stored, routed, and surfaced across every product at planet scale. As a Senior Product Manager for ODP, you will work with world-class engineers and cross-functional partners to shape how the platform is deployed, controlled, and operated. You will define product direction across the control plane and data layer, translate complex infrastructure trade-offs into clear roadmap decisions, and help customers get the most from their observability investment – regardless of architecture, topology, or scale. At Datadog, we place value in our office culture – the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You Will Do: Develop a deep understanding of the Observability Data Platform customers – platform engineers, SREs, and product managers that own the product verticals – their infrastructure challenges, deployment topologies, and cost-to-serve trade-offs. Define product direction across multiple ODP surfaces, including the control plane and data layer, by articulating clear problem statements and desired outcomes, and partnering with engineering on technical approach and sequencing Lead conversations with design partners and strategic customers to understand real-world platform pain points, validate product assumptions, and guide solutions from early prototypes through General Availability Develop a co

AIGoRustSpring
L
📍 New York, NY, United States· Full-time
✓ High-confidence listingCompany trend +33.3%
Quick readStrong listing-quality and freshness signals

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With a billion rides per year and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Rider, Driver, Marketplace, and beyond. While traditional approaches to optimization and problem decomposition are sufficient to disrupt transportation, building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives warrants modern ML utilizing petabyte-scale data. Our highly motivated Machine Learning Engineers work on these challenging problems and define solutions to directly impact various aspects of our core business. The Fulfillment group, within the Marketplace at Lyft, is responsible for determining what inventory can be reliably offered for a given rider session and fulfilling rider requests. The group comprises several sub-teams that generate feasible offers for riders, match rider requests with drivers, and maintain a distributed state machine to track rides and drivers from request through completion. We are seeking a Machine Learning Engineer to join the Fulfillment team and lead the design, development, and deployment of state-of-the-art machine learning systems. This role requires a strategic thinker who can balance high-level system architecture with hands-on technical implementation. You will collaborate across teams to shape the future of ride-sharing by leveraging machine learning and data science. Responsibilities: Design, build, and deploy machine learning models for real-time applications, including translating state-of-the-art research into production-ready solutions Design and implement feature pipelines, model training workflows, and serving infrastructure using Lyft's ML platform Evaluate ML system performance against business KPIs, run experiments, and drive continuous model improvement

PythonMachine LearningAIGo
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio

PythonSQLAWSCI/CD
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. Today, CLEAR is well-known as a leader in digital and biometric identification, reducing friction for our members wherever an ID check is needed. We’re looking for a Senior Software Engineer to establish our Observability framework and foundations. You will join us to accelerate building and scaling our innovative systems that support our growing identity platform. You will drive on Observability best practices to find and fix gaps in our observability and our overall systems. You will also lead practices such as load testing, capacity planning, game days, chaos testing, and incident post-mortems. What You Will Do: Embed within the Engineering pillar to deeply understand the product and implement observability across all key flows Facilitate and build load testing cases, ensuring we understand the limits and scaling factors of our services and systems Contribute to observability and support the design of new services and systems, ensuring highly reliable and scalable concepts are implemented Build and lead practices such as game days, chaos engineering, and failure analysis Build long-term capacity plans, with an eye toward reliability and cost-efficiency Who You Are: 6+ experience writing production-grade software in a modern language, such as Java and Python. Strong knowledge of distributed systems concepts (think CAP theorem), microservices architecture, and distributed tracing . Experience with modern observability systems such as Datadog. Experience with performance debugging tools and patterns. You should be able to read a f

PythonJavaGitRest
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $320K/yr

Quick readStrong listing-quality and freshness signals

As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental research problems and collaborating with Datadog's product and engineering teams to translate research advances into products. Building on our track record of AI-powered solutions (e.g., Bits AI , Bits Evolve , and our time series foundation model ), Datadog AI Research tackles high-risk, high-reward problems grounded in real-world challenges in cloud observability and security. We are focused on two research areas: World Models for Observability -- Training multimodal foundation models that learn the joint dynamics of distributed systems across metrics, traces, logs, topology, and events. These models power advanced forecasting, anomaly detection, root cause analysis, counterfactual simulation ("what if?"), and provide a learned planning backbone for our autonomous agents. Trained Agents for Observability -- Post-training models to operate autonomously across Datadog's domain. SRE incident response is our first target, with a clear path to code repair, security response, and infrastructure optimization. We build the simulation environments, RL training loops, and evaluation infrastructure needed to train agents that match or surpass frontier models at a fraction of the cost. What You'll Do: Conduct research in generative AI and machine learning, building specialized foundation models and trained agents for observability Train multimodal models on large-scale, diverse telemetry data (metrics, logs, traces, topology, events) using distributed training infrastructure Design and build simulated environments and RL training loops for on-policy agent training and evaluation Collaborate with cross-functional teams (Product, Engineering) to integrate capabilities like multimodal world modeling and autonomous agents into Datadog's products Stay at the forefront of foundation models, world models, and RL-based agent research Contribute to r

GitMachine LearningAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $280K/yr

Quick readStrong listing-quality and freshness signals

Datadog is seeking a Director of Product Management for Platforms to lead the internal platforms that power our global engineering organization. This is both a customer and internal platform leadership role, focused on enabling Datadog customers to maximize value with Datadog but also to enable other Datadog products to build, scale, and operate products efficiently and reliably. In this role, you will bring a combination of technical expertise, product management experience, and a deep understanding of platform and shared capabilities to help Datadog grow its leadership position. You will lead a team of product managers and collaborate with senior leadership in product, engineering and design. Your scope includes driving product features shared across all Datadog but also large-scale Datadog’s platform solutions. What You’ll Do Own the vision and strategy for platform products, ensuring alignment with overall company goals and customer needs. Identify new opportunities for innovation and drive them from concept to execution, ensuring they have a measurable impact on customers and the business. Define product roadmaps and manage the prioritization of features and initiatives to ensure the team's efforts are aligned with business goals. Drive Platform Adoption: Partner with engineering and product teams to ensure widespread adoption of shared platforms and shared features, reducing duplication and accelerating delivery. Improve Operational Efficiency: Improve developer velocity, time-to-production, and operational efficiency across Datadog’s engineering ecosystem. Collaborate with cross-functional teams including engineering, design, data science, marketing, and sales to deliver infrastructure product solutions that meet customer needs and business objectives. Define and Track Success Metrics: Define and track platform success metrics, including adoption of platform capabilities, reduction in internal toil, time-to-production improvements, cost efficienc

P
📍 New York, NY, United States
✓ High-confidence listing

$170K – $250K/yr

Quick readStrong listing-quality and freshness signals

A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Optimize cloud financial operations to maximize value from cloud investments, including rapidly growing artificial intelligence (AI) and machine learning workloads Provide actionable insights on cloud spend, SaaS license optimization, and emerging AI cost drivers, including model inference and usage-based consumption Implement tooling, tagging standards, and processes that improve cost visibility and optimization across cloud, SaaS, and AI workloads Monitor large language model API consumption and GPU-intensive infrastructure to identify cost trends, anomalies, and optimization opportunities Build financial models to forecast cloud, SaaS, and AI expenditures for budgeting cycles, commitment decisions, and vendor negotiations Design cost allocation, tagging, showback, and chargeback models that attribute spend to the teams, applications, and use cases driving it Educate engineering and business owners on cloud financial management practices th

AWSAzureMachine LearningArtificial Intelligence
LA
📍 New York, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

SUMMARY STATEMENT We are looking for a Solution Architect to design the technical solutions behind our client engagements and give delivery teams a clear, workable path from concept to production. You will work across enterprise data, software applications and GenAI - translating complex business problems into practical architectures that delivery teams can build and scale. This could include architecting an agentic workflow for clinical operations, a conversational analytics product grounded in enterprise data, or an AI-enabled decision platform for commercial teams. You will work directly with clients, define the architecture, test the most important technical decisions yourself and establish the foundations for successful delivery. This is an architecture-first role with meaningful hands-on engineering: you will stay close enough to implementation to prove the architecture works and support it through production delivery, without becoming the primary engineer for every component. You will also help shape the reusable patterns, technical standards and accelerators behind Lynx’s growing AI-native life sciences practice. KEY RESPONSIBILITIES Solution Architecture Own the end-to-end solution architecture for client engagements, including data models, system design, integration patterns and technology choices. Translate business requirements into clear technical designs and implementation paths that delivery teams can build from. Design solutions spanning enterprise data, APIs, applications, cloud platforms and GenAI capabilities. Lead technical discovery with clients: understand requirements, assess existing systems and identify dependencies, constraints and delivery risks. Present architectural options and trade-offs clearly to technical teams, business stakeholders and senior leaders. Make pragmatic decisions across build speed, cost, scalability, security and maintainability. Review key implementation decisions and remain

TypeScriptPythonAWSAzure
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $156K/yr

Quick readStrong listing-quality and freshness signals

As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to

GitAIRustTerraform
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $276K/yr

Quick readStrong listing-quality and freshness signals

Team description At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the platform that connects these agents to Datadog: the MCP Server, the tools and retrieval surfaces agents call into, and — critically — the evaluation systems that tell us whether an agent's experience on Datadog data is actually getting better over time. This role is about that last piece. We're hiring a Staff Applied Scientist to define what "good" means for an Agentic interface at Datadog and to build the measurement systems that make it true. "Good" isn't one number — it spans answer quality, tool-selection accuracy, retrieval relevance, latency, token cost, and end-to-end agent success on real customer workflows. You'll design the evals, build the datasets, define the metrics, and partner with the AI engineers on the team to land the platform that lets every product group at Datadog ship integrations that are demonstrably better release over release. The space is full of open research questions. How do you evaluate an agent end-to-end when the trajectory is non-deterministic? How do you score tool selection when the tool catalog has hundreds of entries and grows weekly? How do you build a measurement system that catches regressions across first-party and third-party agents at once, without each team writing their own harness? If those are the problems you want to spend your time on, come build this with us. Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply. What You’ll Do: Own the evaluation strategy for Datadog's AI agent integrations. Define the metrics — offline and online, quali

AIGoRustSpring
M
📍 New York, new york, United States· Full-time
✓ High-confidence listingCompany trend -67.9%
Quick readStrong listing-quality and freshness signals

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. About Modal Design Modal is building the future of serverless computing, and the brand that carries that story is still taking shape — You'll join Modal's newly formed Brand team inside our design org as one of its first senior hires, working directly with the Director of Brand Design and Head of Design to build a brand developers recognize instantly and remember. The Role As a staff-level Brand Designer, you will have major influence over every brand surface: the marketing website, campaigns, events, editorial projects like the GPU Glossary and forthcoming publications, and out-of-home work as we scale into larger formats. You'll also be a beacon to external agencies, representing Modal's internal creative voice and making sure the work translates into a system we can actually build on. And as the studio grows, you'll help set its craft standard — guiding and mentoring earl

GitAIGoMarketing
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc

SQLKubernetesGitLinux
C
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

From £270K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Role Overview: As a Senior Research Engineer in our Safety team, you will play a key role in helping develop safer, more secure, and more reliable models. Your primary focus will be on building tools to enable easy data synthesis, analysis, and management, for complex combinations of real and synthetic data that is used in both model training and evaluation. You will own the cohesive vision of these tooling repositories. You will work closely with a team of research scientists and engineers to create tooling that enables tighter experimentation cycles, better data coverage of the real world, and more scientific rigour. You will have a lot of autonomy and need to be opinionated about what areas of the codebase need elegance and standards, and where that would be overengineering. You will be given high level experimental problems that need to be solved with efficient pipelines, and design and implement the solutions. Your data analysis will collaboratively feed into modelling decisions and experimentation. This role combines expertise in software engineering, statistics, and data science. If any of these topics sound interesting t

PythonSQLGitRest
🔔

Get new design and cost estimation head jobs in New York, United States by email

Daily job updates · Unsubscribe anytime