About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
Jobs in United States
Customer Success Engineer Devops in New York
449 active opportunities · Updated October 2026
Showing
15 jobs
Explore current customer success engineer devops jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already run multi-node training, elastic inference, sandboxes, and distributed volumes, and we control the infrastructure underneath. We’re looking for research depth in post-training to sit alongside our systems and product work. What you'll do: We are looking for research scientists with a strong track record in reinforcement learning, machine learning, and foundation models, including large language and multimodal models, to join our research team. This role is well suited to candidates interested in improving existing methods and developing new techniques for large-scale model training, optimization, and inference, extending models to long-context and long-horizon tasks, and improving inference-time efficiency, reliability, and robustnes
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads. As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, e
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit
From £270K/yr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Role Overview: As a Senior Research Engineer in our Safety team, you will play a key role in helping develop safer, more secure, and more reliable models. Your primary focus will be on building tools to enable easy data synthesis, analysis, and management, for complex combinations of real and synthetic data that is used in both model training and evaluation. You will own the cohesive vision of these tooling repositories. You will work closely with a team of research scientists and engineers to create tooling that enables tighter experimentation cycles, better data coverage of the real world, and more scientific rigour. You will have a lot of autonomy and need to be opinionated about what areas of the codebase need elegance and standards, and where that would be overengineering. You will be given high level experimental problems that need to be solved with efficient pipelines, and design and implement the solutions. Your data analysis will collaboratively feed into modelling decisions and experimentation. This role combines expertise in software engineering, statistics, and data science. If any of these topics sound interesting t
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the data infrastructure behind some of the most demanding AI training workloads in the world, and we want sharp, curious people to help us do it. In this role, you'll build and maintain the high-performance data layer our Modeling teams rely on for training and evaluation jobs. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it. Collaborate daily with researchers and engineers who are some of the best in the world at what they do. You may be a good fit if you have: 4+ years of experience working on data storage infrastructure Strong command of Python Kubernetes experience, especially on the storage side (Persistent Volumes, CSI drivers, etc.) The ability to transform unstructured data into performant datasets across diverse storage backends including S3, GCS, and POSIX Experience with distributed data processing frameworks such as Apache Beam, Spark, or Flink [Nice-to-have] Familiarity with modern analytics tooling such as BigQuery, Airflow, or dbt Genuine excitement about AI.
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Our Sales team combines deep product and industry knowledge and is focused on bringing Plaid to an ever-broadening set of businesses. Our thesis is that every company in financial services, and most specifically the Banks themselves, can benefit from better technology and access to data. We’re building a team to work exclusively with the financial institutions throughout North America to help enable their transformation to deliver incredible digital and consumer experiences to their customers, at the scale and quality needed by the largest banks. Your focus will be on selling into a territory of banks and credit unions, running the full sales cycle and building strong, long-lasting relationships to help them execute on their vision. You'll bring Plaid to banks and credit unions, with a focus on banking and wealth workflows like online account opening, consumer lending, asset and income verification, and fraud prevention. Responsibilities: Identify potential customers and run the end-to-end sales process within our Banking and Wealth segment Build and maintain relationships with Financial Institutions, from executives to product teams and developers; be viewed as a subject matter expert Partner with
From $252K/yr
We're looking for an Engineering Manager II to own and grow the Observability Pipelines engineering org at a pivotal moment in the product's lifecycle. Observability Pipelines is Datadog's on-premise, vendor-agnostic telemetry pipeline product, with a lot still to build as it grows and scales. It sits at the center of a fast-consolidating market, is central to Datadog's data pipeline optimization story for Logs and Metrics customers. This is a build-and-scale opportunity: you'll grow the management and technical leadership layers, co-own the roadmap with Product, and define how this org operates as it continues to expand. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You'll Do: Directly manage the OP org including EM1s across NYC and Paris, set technical direction, and be the connective tissue across a distributed team Build out the management and technical leadership layers as the org continues to grow - today ~20 ICs Partner directly with Product to co-own the roadmap and strategy, helping decide where OP’s engineering investment goes next Set and evolve the operating rhythm across the group: planning cadence, on-call and incident standards, and cross-team alignment Own key cross-org relationships with the SaaS Logs Pipelines team, the BYOC team, and the Vector open-source community Coach managers and senior engineers, and build the succession and growth plans that let the org scale beyond you Who You Are: Experienced managing managers across distributed teams, with a track record of raising the bar on how those teams operate, not just delivering through them Back
From $192K/yr
Datadog's Application Performance Monitoring (APM) provides deep visibility into the health, performance, and lifecycle of modern distributed applications, tracing requests from end-user devices (web and mobile) through to backend services. Our goal is to help customers detect root causes faster, optimize application performance, and improve resource efficiency at scale. As the Engineering Manager for APM Serverless, you will help define and deliver the end-to-end serverless APM experience, from auto-instrumentation through troubleshooting, and ensure that OpenTelemetry and Datadog-native customers alike have a frictionless and performant journey. You will also lead efforts to expand coverage of cloud-managed services across providers, ensuring customers can seamlessly trace and monitor critical services in all major and emerging cloud environments. We’re looking for an experienced engineering leader who thrives at the intersection of infrastructure and developer experience. You should care about well-designed APIs, observability-first thinking, and building systems that empower other developers. This is a high-leverage role that will influence how developers across the industry understand and instrument their serverless workloads. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Lead a polyglot team of 8-9 engineers and partner closely with Product and Engineering teams across Datadog to deliver industry-leading serverless capabilities that power consistent, scalable, and intuitive instrumentation across languages. Drive a domain that is technically rich: Lambda, Azure Functions, GCP, OTel billing, Rust, durable functions, distributed tracing across managed services. Engineers on this team work
From $192K/yr
Databases and data stores are at the center of our applications, for legacy applications and modern AI applications alike. Yet most observability and optimization approaches lack a holistic approach or application context. Datadog has been on a mission to revolutionize how databases are operated, flipping what is often seen as a black box of complexity prone to security and performance risks, into a well oiled machine enabling our builders and businesses to move faster and smarter. We’re looking for an experienced product manager passionate about joining this mission to lead this product opportunity. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead ambitious investments that rethink how customers operate and get value from their databases, diving into ambiguity, working with our customers on new models, and shipping new products and changes to existing products. Develop a deep understanding of our customers and their issues, what problems are really behind those issues, and how we can improve how databases are operationalized across SRE teams, DBAs and application developers. Continuously refine your understanding of database management systems and datastores from SQL and OLTP based to NoSQL e.g. AWS RDS, PostgreSQL, SQL Server, MongoDB, MySQL, etc Define, build and launch the next generation of database monitoring and optimization capabilities for our customers Join a talented engineering team with a record of disrupting observability approaches to further the mission of demystifying and optimizing databases using your team’s creativity, alongside your customers’ problems, as a key resource. Collaborate with other Product teams in Datadog to maintain and improve all Datadog products, improve the seamless integration across th
Role Overview We’re looking for a Senior Product Designer with a passion for building platform features that impact UX at a massive scale. As a member of AAA, you’ll work across a breadth of problem spaces on deeply technical topics, such as governance, access, teams, identity, and multi-org UX. You’ll lead design for features that mediate how Datadog’s biggest customers use the platform and manage what their employees can see and do. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do AAA is a highly technical product environment, where designers strive to understand the technical workflows, constraints and opportunities within this domain. As you develop subject matter expertise, you’ll use your influence as a Senior designer to drive the roadmap and vision. Day-to-day, you’ll work closely with your PM and engineering partners to understand customers’ pain points, explore solutions, and build UI that ships. Prototyping in code creates the opportunity to share high-fidelity ideas at early stages, even if the design itself is low-fidelity. Tight feedback loops let your team iterate quickly, building towards an ideal vision for the best possible experience while you ship incremental steps towards that goal. In addition to making the work, you will also share your work with stakeholders at varying levels of seniority, all the way up to executive leadership level. Your crisp context setting and clear description of your solution will help people understand your choices and sell them on your proposal. Your work will up-level the quality bar for what an admin experience can look like. While you’ll leverage Datadog’s design system where it makes sense, you won’t let it prevent broad exploration of different design patterns in the service of finding the optimal solu
From $71K/yr
Datadog is looking for a resourceful and creative Associate Field Marketing Manager to lead event strategy and execution for Datadog's AI product line across the East and Canada region. This role is ideal for someone who is passionate about AI and developer communities, enjoys getting hands-on with technical audiences, and wants to build a market-leading brand presence for Datadog's AI offerings. As part of the NAMER Field Marketing team, you will own the strategy, planning, and execution of a mix of practitioner-focused events, hands-on workshops, and surround activations at major AI conferences. This role is critical to scaling awareness and adoption of Datadog's AI products, and to building durable relationships with AI customers, prospects, and the broader developer community. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Own the strategy, planning, and end-to-end execution of AI practitioner events, hands-on workshops, and meetups across the East and Canada Lead surround and off-site activations at major AI conferences (e.g., NVIDIA GTC, Ray Summit, AI Engineer Summit, and similar industry events) to build brand visibility and drive engagement with target audiences Partner with Product Marketing and AI/ML product teams to translate Datadog's AI observability and LLM monitoring capabilities into compelling, technically credible event content Design and continuously improve hands-on workshop curriculum and live demos that showcase Datadog's AI products to practitioners and technical decision-makers Build scalable, repeatable event playbooks and toolkits so programs can be run consistently across multiple markets Manage vendors, venues, budgets, staffing, and on-site logistics, ensuring every event reflects Datadog's brand and delivers a seamles
About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf
$140K – $175K/yr
About The Weather Company: The Weather Company is the world’s leading weather provider, helping people and businesses make more informed decisions and take action in the face of weather. Together with advanced technology and AI, The Weather Company’s high-volume weather data, insights, advertising, and media solutions across the open web help people, businesses, and brands around the world prepare for and harness the power of weather in a scalable, privacy-forward way. The world’s most accurate forecaster globally, the company reaches hundreds of enterprise clients and more than 360 million monthly active users via its digital properties from The Weather Channel (weather.com) and Weather Underground (wunderground.com). AI is part of how we work: Human judgment, expertise, and creativity remain essential and are amplified by AI. We expect everyone in this role to use AI thoughtfully and proactively to accelerate their performance, improve the quality of their work, and explore new possibilities. We are looking for people who are curious, adaptable, and excited to help shape an AI-enabled culture that delivers better outcomes for our customers and our company. Job brief: We are seeking a Lead of Data Instrumentation & Growth Measurement to serve as the connective tissue between Product (Web, iOS, Android) and Marketing (Lifecycle, Performance Marketing). In this role, you will eliminate data disparity, establish a unified enterprise event taxonomy, manage Server-Driven UI (SDUI) dynamic tracking schemes, and lead incrementality testing to ensure marketing spend is optimized. You will also oversee performance marketing capabilities, configuring Server-Side Tagging (CAPI), Google Tag Manager (GTM), and a targeted GA4 conversion infrastructure. You will partner with product managers, developers, and specialized agency partners to overhaul our Amplitude environment and build transparent attribution pipelines. The impact you'll
Other cities to consider
More places hiring for this role
Get new customer success engineer devops jobs in New York, United States by email
Daily job updates · Unsubscribe anytime