Jobs in United States

Model Behavior Engineer in New York

171 active opportunities · Updated October 2026

Explore current model behavior engineer jobs in New York. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 New York, New York, United States· Full-time
✓ High-confidence listingExact matchCompany trend -86%

$98K – $140K/yr

Quick readExact title match for your search

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role You'll own the quality bar for Notion AI products. You’ll work with product and engineering teams to build systems to define what “good” looks like, measure our progress, and drive changes to deliver reliable and high-quality AI experiences. Your work directly shapes how Notion's AI products behave for millions of users. This isn't a traditional software engineering role. It’s an art & science role . You won't spend your days writing code. Instead, you'll focus on understanding and shaping how our AI products behave through context engineering, designing evaluation systems, and analyzing data. This team sits in our AI engineering team, working directly with engineering, product, design, and data. This role is a unique blend of ops, strategy, and product thinking. Day to day, you'll live in production data, ship prompt fixes, run evals and, in effect, shape our quality strategy. As part of that you'll shape Notion's model strategy and work directly with frontier AI labs (OpenAI, Anthropic, Google) to evaluate and launch new models. We're looking for problem-seeking generalists interested in 0 → 1 : curious people wi

SQLRestAIGo
O
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team API Frontiers turns OpenAI’s frontier models into production APIs that developers can use to build reliable products and agents. We own the core path connecting models to developers through the Responses API, with a focus on safety, reliability, and speed. Working closely with Research, Safety, Codex, and other API teams, we bring new model capabilities into production and improve them through developer feedback. About the Role We are looking for a backend software engineer to build and operate the services behind the Responses API. You will shape API behavior, bring new capabilities from research into production, and make long-running agent workflows dependable and fast. The work combines distributed systems engineering with product judgment: designing useful developer interfaces, managing staged rollouts, and following production issues through to durable fixes. In this role, you will: Design, build, and operate APIs and backend services that bring frontier model capabilities to developers. Partner with Research, Safety, Codex, and API teams to define API behavior and deliver safe, staged launches. Build API capabilities for agent workflows, including task delegation, context sharing, and parallel execution. Strengthen long-running request reliability across timeouts, cancellation, streaming, and background execution. Improve request-processing performance and tail latency through profiling, efficient systems code, and persistent connections. Turn developer feedback and production failures into better observability, diagnostics, and lasting product improvements. Your background might look something like: 5+ years of experience building and operating backend services or developer-facing APIs in production. Strong software engineering fundamentals, with practical knowledge of distributed systems, concurrency, and asynchronous execution. Ability to diagnose production failures and performance bottlenecks using observability data and profiling. Product

AWSRestAIRust
O
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team OpenAI’s Forward Deployed Engineering team partners with customers to turn research breakthroughs into production systems. We operate at the intersection of customer delivery and core platform development. About the role As a foundational FDE manager, you’ll lead FDE through high-stakes, ambiguous customer deployments and own technical and business value outcomes end to end. You’ll grow a team that can operate under pressure and help OpenAI learn from the field. You’ll partner closely with Product, Research, Sales, and GTM to ensure fieldwork informs roadmap priorities, drives new exploration, and supports safe deployment at scale. Your decisions will influence how OpenAI is trusted by the customers closest to our deployment work. Your success will be measured by how consistently your team ships, how clearly you deliver signal to Research and Product, and how durable your team and delivery model prove to be. This role is based in New York City We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. This role also will require travel up to 25%. In this role you will Lead and grow a team of FDE delivering production systems with frontier models Own end-to-end delivery outcomes through clarity, speed, tight coordination, and technical quality Codify what works into tools, playbooks, and roadmap inputs that create leverage for both OpenAI and our wider developer community Notice early indicators and raise them with urgency, whether in product behavior, customer environments, or delivery practices Use judgement to distinguish what requires action and what does not Set a high bar for FDE performance and support each person’s growth through direct, actionable feedback Define how we staff and support field teams that can scale without added complexity You might thrive in this role if you Bring 8+ years of engineering or technical delivery experience, including 2+ years managing high-performing FDE or custo

JavaScriptPythonJavaAWS
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect

PythonAWSMachine LearningAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As the Senior Product Manager for the Actions & Automations team, you will own the ecosystem that enables customers, partners, and Datadog teams to build, deploy, and operate AI agents on Datadog. You will drive the strategy and execution for the platform capabilities, developer experience, integrations, and extensibility model that make Datadog the best place to build agents that understand and act on production systems. Modern engineering organizations are entering a new era where software is not only monitored and operated by humans, but increasingly by AI-powered agents. As agentic workflows reshape how teams build, operate, secure, and troubleshoot systems, customers need a platform for creating specialized agents, connecting them to business and engineering systems, governing their behavior, and extending them to solve unique organizational problems. You will define and build the ecosystem that makes this possible. Agent Builder sits at the intersection of Datadog's products, AI capabilities, and ecosystem strategy. You will have the opportunity to work across the breadth of the Datadog platform, partner with teams throughout the company, and help establish Datadog as the foundation for operational AI. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define the vision, strategy, and roadmap for Datadog's Agent Builder platform and ecosystem. Own the core platform capabilities that enable customers and partners to create, customize, deploy, and manage AI agents. Drive the extensibility model for agents, including integrations, tools, actions, context sources, APIs, SDKs, and developer workflows. Shape how agents perform actions across Datadog products and third-party systems. Partner closely with AI, platform, infrastructure, and product tea

AIGoRustSpring
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc

SQLKubernetesGitLinux
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit

GitRestMachine LearningAI
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may

PythonGitRestMachine Learning
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads. As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, e

PythonGitRestAI
H
📍 New York, NY, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Most AI engineering jobs are a thin wrapper around a model API. This role is different. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to the source document, and route ambiguous cases to human experts for review. Our users make decisions that impact real healthcare outcomes, so “good enough” is not good enough. Building AI systems that are accurate, reliable, auditable, and scalable is at the core of this role. As a Senior AI Applied Engineer, you will design, build, deploy, and operate production AI systems used at scale within one of the largest health insurers in the United States. You will own solutions end-to-end, from user experience and APIs to model orchestration, evaluation frameworks, infrastructure, and production operations. Why Join Us Build production AI systems where LLMs are in the critical path, not just demos or proofs of concept. Work on extraction, retrieval, agentic workflows, and human-review systems that process real healthcare data at scale. Own projects end-to-end across frontend, backend, AI orchestration, infrastructure, deployment, and operations. Solve challenging problems around accuracy, explainability, traceability, and reliability in regulated environments. Ship quickly in a small, high-impact team that embraces AI-assisted development and rigorous quality standards. Build systems that continuously improve through expert feedback, evaluations, and human-in-the-loop workflows. Key Responsibilities Design, develop, and deploy full-stack AI-powered application

JavaScriptTypeScriptPythonReact
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed ML Engineers who want to work at the intersection of deep technical work and direct customer impact. As an ML FDE, you'll partner with leading AI companies and foundation model labs to help them achieve state-of-the-art performance on their most demanding workloads — LLM serving, model training (SFT, RLHF), audio pipelines, scientific computing, and more. You're helping teams reach outcomes most engineers can't on their own. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the AI stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and optimize production AI workloads

RestAIGoRust
C
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

From £270K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Role Overview: As a Senior Research Engineer in our Safety team, you will play a key role in helping develop safer, more secure, and more reliable models. Your primary focus will be on building tools to enable easy data synthesis, analysis, and management, for complex combinations of real and synthetic data that is used in both model training and evaluation. You will own the cohesive vision of these tooling repositories. You will work closely with a team of research scientists and engineers to create tooling that enables tighter experimentation cycles, better data coverage of the real world, and more scientific rigour. You will have a lot of autonomy and need to be opinionated about what areas of the codebase need elegance and standards, and where that would be overengineering. You will be given high level experimental problems that need to be solved with efficient pipelines, and design and implement the solutions. Your data analysis will collaboratively feed into modelling decisions and experimentation. This role combines expertise in software engineering, statistics, and data science. If any of these topics sound interesting t

PythonSQLGitRest
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $280K/yr

Quick readStrong listing-quality and freshness signals

Datadog is seeking a Director of Product Management to lead our AI Observability portfolio and shape how organizations build, monitor, and scale AI systems in production. This role leads LLM Observability and helps define the next wave of innovation across GPU Monitoring, Distributed AI Monitoring, and emerging research-oriented tooling such as Model Lab. You will set the vision and strategy for this rapidly growing area, expanding established products while incubating new capabilities that deliver deep visibility into AI infrastructure, model performance, and distributed AI environments. As AI becomes core to modern applications, this team plays a critical role in ensuring customers can deploy and scale AI with confidence. We’re looking for a builder-minded product leader with strong technical depth and hands-on curiosity - someone who has built or worked closely with AI-powered products and understands the realities of production AI. You will lead a team of product managers and partner closely with engineering and design to advance Datadog’s leadership in AI observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the vision and strategy for AI-driven products, ensuring alignment with overall company goals and customer needs. This will include managing our embed program to enhance the capabilities of existing products as well as developing dedicated and independent AI products. Lead and mentor a team of product managers, helping them grow and advance their careers while ensuring the delivery of high-quality, AI-powered features. Collaborate with cross-functional teams including engineering, data science, marketing, and sales to deliver AI product solutions that meet customer needs and business objectives. Identify new opportunities for

Machine LearningAIGoRust
C-
📍 New York, New York, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. CLEAR1 is CLEAR's B2B identity verification business, serving enterprise, healthcare, and workforce customers through direct sales and partners. We are growing fast and need a Revenue Operations Leader who can build the operating model to match. This is a strategy and accountability role. You will own how the go-to-market team plans, measures, forecasts, and runs its process across the full customer lifecycle, from first touch through renewal and expansion. You will define what our revenue numbers mean and where they come from, make them consistent from rep to Board, and hold the team to the process that produces them. You will be the CRO's partner on where to invest and what is working. You have done this before at a growth-stage B2B company, you think like a finance partner as much as an operator, and you direct systems work without needing to do it yourself. What you'll do: Own CLEAR1’s revenue operating model, partnering with the CRO and Finance on annual and quarterly planning across capacity, quotas, coverage, pipeline targets, and investment decisions Establish a single source of truth for revenue performance by defining core metrics, data sources, and reporting across ARR, net new ARR, GRR, NRR, pipeline coverage, win rate, sales cycle, and CAC payback; own executive and Board-level reporting Lead forecasting and revenue accountability end to end across new business, renewals, and expansion, establishing clear definitions, cadence, and ownership while continuously improving forecast accuracy Build and enforce scalable processes

AISalesforceFinanceCRM
O
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the team OpenAI’s Forward Deployed Engineering team partners with banks, asset managers, insurers, and private capital firms to deploy production-grade AI systems in high-stakes financial environments. We operate at the intersection of customer delivery and core platform development, embedding deeply with customers to translate frontier model capabilities into reliable, auditable systems that create measurable business impact. Our work turns early deployments into repeatable solution patterns, operating standards, and evaluation practices that scale across regulated financial institutions. About the role We are hiring a Forward Deployed Engineer (FDE) to lead end-to-end deployments of OpenAI’s models inside financial services organizations where correctness, latency, explainability, and control matter. You will work with customers who are experts in investment banking, trading, risk, compliance, underwriting, research, operations, or investment decision-making, translating complex workflows, data constraints, and regulatory requirements into production systems. You will measure success through production adoption, workflow efficiency, risk reduction, revenue impact, and evaluation-driven feedback loops that inform product, model, and GTM strategy. You’ll work closely with Product, Research, GTM, Security, Legal, and GRC to deliver systems that meet enterprise standards for governance, auditability, and operational resilience. You will also play a central role in shaping OpenAI’s Financial Services offering — identifying high-value use cases, defining solution patterns, and building the first repeatable deployments that scale across institutions. Learn more about some of our work with financial institutions . This role is based in New York. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. Travel up to 50% may be required. In this role, you will Design and ship production AI systems around models, owning integrations,

JavaScriptPythonJavaAWS
🔔

Get new model behavior engineer jobs in New York, United States by email

Daily job updates · Unsubscribe anytime