Jobs in United States

Model Behavior Engineer in United States

2,174 active opportunities · Updated October 2026

Explore current model behavior engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We're hiring a hands-on Engineering Manager to build and lead Replit's Anti-Abuse team from the ground up. This is a foundational 0-to-1 role: you'll define the anti-abuse roadmap, hire a small team of engineers and data analysts, and ship the systems that protect Replit's platform, users, and economics from adversarial actors. You'll partner across Support, Legal, Security, Infrastructure, and the Money and Growth teams to make abuse economically unviable while keeping friction low for legitimate users. Replit sits at the frontier of AI-native abuse. Our platform is a target for phishing and scam hosting, cryptomining, LLM token farming, card and coupon fraud, and increasingly, abuse driven by AI agents themselves. The team you build will define how Replit defends against all of it. What You'll Do Build the anti-abuse roadmap from scratch : Define the threat model, prioritize across abuse vectors (phishing/scam hosting, cryptomining, token farming, payment fraud, AI agent exploitation), and translate it into a shipping plan with clear sequencing and tradeoffs. Design progressive verification and identity infrastructure : Build the "ladder of trust" that gates increasing platform capabilities (referrals, additional credits, access to powerful agent features, Missions) behind escalating verification. This includes a humanity/identity layer that's distinct from user accounts, integrations with KYC-grade verification providers, and the policy engine that decides what level of trust unlocks what behavior. This infrastructure is core not just to promo integrity but to how Replit safely expands agent capabilities over time. Ship as a hands-on EM : Stay in the code. Use the latest AI coding tools (including Rep

GitAIGoRust
R
📍 Foster City, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.9%
Quick readStrong listing-quality and freshness signals

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for an AI Agent Security Architect to function as the primary technical authority for Replit’s autonomous and AI agent security blueprint. In this critical role, you will design, implement, and maintain the runtime defense systems, guardrail frameworks, and sandboxing architectures that govern AI agents executing code, invoking tools, and reasoning across our platform. You will be a key technical contributor—leading high-impact AI security initiatives and bridging the gap between non-deterministic AI behavior and rigorous cybersecurity controls for both engineering and executive leadership. What You'll Do AI Agent Security Strategy & Technical Execution AI Agent Security Blueprint: Define the long-term vision and architectural patterns for securing autonomous agent workflows, Model Context Protocol (MCP) integrations, multi-turn reasoning loops, and multi-agent coordination. Runtime Guardrails & Policy Enforcement: Architect and deploy dynamic input/output guardrail systems, semantic firewalls, and real-time intent verification filters to prevent goal hijacking, system prompt leaks, and indirect prompt injections. Agent Execution & Tool Sandboxing: Partner with Infrastructure and AppSec teams to design secure, short-lived, micro-isolated environments (e.g., microVMs, WebAssembly, container sandboxes) where agents can dynamically execute code, run shell commands, and interact with host operating systems safely. Agentic Threat Modeling & Red Teaming: Conduct specialized threat modeling against non-deterministic systems. Lead automated and manual AI red-teaming initiatives to uncover vulnerabilities in RAG context pipelines, vector stores, and tool-calling interfaces. Identity

JavaScriptTypeScriptPythonJava
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team: The OpenAI API team builds the foundation that enables every developer to harness OpenAI’s models safely, reliably, and at scale. We design and operate the systems that power model serving, API access, billing, developer tooling, and enterprise integrations—forming the connective tissue between OpenAI’s research breakthroughs and real-world products. Our mission is to make it effortless for anyone to build with OpenAI technology. We’re responsible for the infrastructure and product layers that allow millions of developers to integrate GPT models, fine-tune behavior, manage data, and deliver transformative experiences to their users. We collaborate across product, research, and engineering teams to ensure that innovation in model capabilities translates directly into value for customers. The API team spans multiple disciplines, including product management, infrastructure engineering, developer experience, and data systems. We care deeply about reliability, scalability, and simplicity—creating tools that let developers focus on their ideas while we handle the complexity of running world-class AI systems. About the Role: We are seeking an experienced Product Manager to define and scale the construction of our data processing, data privacy, billing, and access controls products. You will set strategy and execute on projects like expanding our regional data processing footprint, enabling new inference caching controls in the API or building APIs that make it easier for organizations to manage their spend limits. You will also define the strategy and ship foundational capabilities that ensure customers use OpenAI products securely, privately, and with enterprise-grade controls. This role partners deeply with engineering, security, legal, compliance, finance and leadership to deliver high-trust, enterprise-grade systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to n

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to train nuanced safety behaviors, how to make the model robust to bad actors, how to address privacy and security risks, and how to make the model trustworthy in safety-critical situations. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role We’re seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You’ll advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. In this role, you will: Research and implement methods for safety training, reinforcement learning, and adversarial robustness. Develop evaluations, identify model failure modes, and use findings to improve training. Work with research, engineering, security, and policy partners to support safe, reliable deployment. You might thrive in this role if you: Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness. Have a degree in computer science, machine learning, or a related field, and strong deep learning research or engineering skills. Have experience improving model safety for deployment and enjoy collaborative research. Are motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings. Security Requirements Active TS/SCI clearance or equivalent. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefi

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin

AWSRestAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As the Senior Product Manager for the Actions & Automations team, you will own the ecosystem that enables customers, partners, and Datadog teams to build, deploy, and operate AI agents on Datadog. You will drive the strategy and execution for the platform capabilities, developer experience, integrations, and extensibility model that make Datadog the best place to build agents that understand and act on production systems. Modern engineering organizations are entering a new era where software is not only monitored and operated by humans, but increasingly by AI-powered agents. As agentic workflows reshape how teams build, operate, secure, and troubleshoot systems, customers need a platform for creating specialized agents, connecting them to business and engineering systems, governing their behavior, and extending them to solve unique organizational problems. You will define and build the ecosystem that makes this possible. Agent Builder sits at the intersection of Datadog's products, AI capabilities, and ecosystem strategy. You will have the opportunity to work across the breadth of the Datadog platform, partner with teams throughout the company, and help establish Datadog as the foundation for operational AI. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define the vision, strategy, and roadmap for Datadog's Agent Builder platform and ecosystem. Own the core platform capabilities that enable customers and partners to create, customize, deploy, and manage AI agents. Drive the extensibility model for agents, including integrations, tools, actions, context sources, APIs, SDKs, and developer workflows. Shape how agents perform actions across Datadog products and third-party systems. Partner closely with AI, platform, infrastructure, and product tea

AIGoRustSpring
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team: The OpenAI API team builds the foundation that enables every developer to harness OpenAI’s models safely, reliably, and at scale. Our mission is to make it effortless for any developer to build transformative products with OpenAI’s models. We’re responsible for the infrastructure and product layers that allow millions of developers to integrate our models, fine-tune behavior, manage data, and deliver experiences to their users. We collaborate deeply across product, research, and engineering teams to drive innovation at the model layer and then directly translate that into value for customers. About the Role: As the Agents Product Manager for the API team, you'll be at the forefront of defining and guiding the future of how developers build agentic applications on top of our AI models. You’ll set clear priorities and drive impactful improvements to model capabilities, balancing user needs, safety considerations, and technical innovation. This role is perfect for a proactive, technically adept PM who thrives on solving challenging, ambiguous problems through structured product thinking and close collaboration with customers, engineers, and researchers. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Deeply understand problems faced by agent builders and identify opportunities where our products and models can make building agents faster, more intuitive, more reliable, and more powerful. Define strategic priorities and roadmap for improving agentic infrastructure for API users, focusing on user outcomes and emerging capabilities. Partner with research and engineering teams at a technical level to translate those priorities into developer products and features (SDKs, APIs, and more). Deliver quickly while maintaining a high bar for product quality and user experience. You might thrive in this role if you: Have 5+ years of product management or related industry experience. Proven track record of b

AWSRestAIRust
P
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About the Team At Embedded Insights, we find the best machine learning opportunities for external products and internal systems, and collaborate with cross-functional partners to bring them to life. We are a central team of Machine Learning Engineers and Data Scientists. We embed with partner teams to build and apply machine learning models that improve internal decision-making and power the Plaid product suite. About the Role You will be the first Data Scientist on the Embedded Insights team, part of Plaid’s Data organization. You will establish the analytics and metrics backbone for a team supporting a diverse set of internal and external products. You will help drive better decision-making, support machine learning model development, and contribute directly to the health of the Plaid network and the quality of Plaid’s products. Your day-to-day work will include: Analyzing entities across the Plaid network to understand behavior and identify opportunities, anomalies, and risks. Creating foundational metrics, dashboards, and monitoring systems that provide a clear view of network health and machine learning model performance. Evaluating the value and performance of machine learni

PythonSQLAWSMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Education team is building products that advance how people learn with AI. The team works across higher education institutions, K-12 districts, and country-level partnerships, including applied research on how AI affects learning and cognitive outcomes. The team owns owns ChatGPT Edu, ChatGPT for Teachers, and related product/research work. The team partners closely with go-to-market, research, Consumer Learning, and model teams to turn education-specific insights into product experiences that can improve ChatGPT more broadly. Some of our recent work: New Education Plugins for ChatGPT Work and Codex New tools for understanding AI and learning outcomes Education for countries Advancements in higher education Early product work - Introducing Study Mode About the Role We’re looking for a hands-on Tech Lead Manager to lead and manage a team of senior full-stack engineers building AI-native learning experiences in ChatGPT. This person will combine technical execution, product judgment, and people leadership: they will write and ship code, manage engineers, and help shape the product direction for how students and Educators use AI. In This Role, You Will Lead and manage a team of three senior full-stack engineers. Build product experiences for ChatGPT Education, ChatGPT for Teachers, and AI-native learning workflows. Partner with research teams on field studies, randomized control trials, classifiers, data pipelines, and cognitive-outcome measurement. Collaborate with Consumer Learning and model teams to translate education insights into broader ChatGPT behavior and product improvements. Drive execution across product, engineering, research, go-to-market, and partner teams. Help define product strategy, priorities, and delivery plans for a new product pod. You Might Thrive In This Role If You Have several years of direct people-management experience with engineers. Are still highly technical and comfortable doing IC engineering work. Have strong pr

AWSRestAIGo
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $1.3M/yr

Quick readStrong listing-quality and freshness signals

The Opportunity MongoDB’s partner ecosystem — systems integrators, ISVs, and the major cloud providers — is a strategic growth engine for the business. We’re looking for a Senior Manager, Sales Plays and Offerings to design the joint go-to-market motions that turn partner relationships into pipeline and revenue. This is a highly cross-functional, strategic role: you’ll build the sales plays and joint offerings themselves, partner with Enablement to get the field and our partners ready to sell them, instrument how they perform, and work with the Programs lead to make sure incentives reward the behavior we want to see. You will report directly to the VP of Partner Strategic Operations and act as a connective layer between Partnerships, Sales, Enablement, and Programs — translating ecosystem strategy into repeatable, measurable, field-ready motions. We are looking to speak to candidates who are based anywhere in the US for our hybrid working model. What You'll Do Build sales plays and joint offerings Design and package partner sales plays and joint solution offerings with priority ISVs, SIs, and cloud partners — defining the joint value proposition, target segment, competitive positioning, and playbook for how field and partner sellers execute it Partner with Product Marketing, Solutions Engineering, and partner counterparts to validate technical integration stories and translate them into a compelling, sellable narrative Prioritize which plays to build and scale based on market opportunity, partner readiness, and alignment to MongoDB’s strategic pillars (e.g., AI, migrations, industry verticals) Own the lifecycle of each play from concept through launch, iteration, and eventual retirement or refresh Drive partner and field enablement Partner closely with the Enablement team to translate each sales play into field- and partner-facing assets: pitch decks, battlecards, demo scripts, certification conte

MongoDBAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Revenue team plays a critical role in enabling OpenAI to scale commercial offerings by overseeing billing operations, deal desk, revenue systems, revenue accounting, and controllership. We work cross-functionally with Product, Engineering, Go-To-Market, Tax, Legal, and Technical Accounting to support new monetization strategies, improve operational efficiency, and maintain financial integrity as the business grows. About the Role This senior leader will own key elements of Ads revenue accounting from technical assessment through operational execution. The role will guide accounting for products, pricing, contracts, incentives, credits, refunds, makegoods, international expansion, and new go-to-market motions. It will establish governance and translate approved accounting positions into launch, billing, data, close, reconciliation, and control requirements. Success requires deep technical revenue expertise, strong business partnership, and the ability to build durable 0-to-1 processes in a fast-changing environment. Advertising is a critical and growing monetization vector for OpenAI, and this role will help shape the financial foundations that enable Ads to scale responsibly and transparently. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead technical accounting assessments for Ads products and commercial arrangements, including performance obligations, variable consideration, allocation, principal-versus-agent, collectibility, contract modifications, refunds, incentives, credits, makegoods, and revenue presentation. Own and continuously evolve Ads revenue accounting policies and operating guidance as product behavior, pricing, contracting, incentive programs, and billing models change. Establish governance for new pro

AWSGitRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight : Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review ). About the Role This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review ). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected. We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader syste

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th

AWSRestMachine LearningAI
B
📍 Colorado Springs, United States
✓ High-confidence listingCompany trend +515.8%
Quick readStrong listing-quality and freshness signals

Lead Systems Engineer - Phantom Works Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Lead Systems Engineer (Level 4) to support the Space Battle Management, Command and Control (SBMC2) programs in Colorado Springs, CO or Berkeley, MO . The Systems Engineer will perform as part of a high-performing team and have the opportunity to contribute to the development and delivery of innovative software solutions in an agile software development environment interfacing with internal and external stakeholders; to include a Joint Industry partnership Team (JIPT) and Working Groups, contributing to the design and development of next generation ground command and control capabilities. As a member of the Boeing team, you will be responsible for key portions of our development lifecycle, from idea creation and development, all the way through to maintenance and support of the customer’s delivered system. More importantly, you will have the opportunity to make an impact on the results of our projects. We offer a collaborative mentoring environment where you have the opportunity to learn from others and be a mentor to others. Position Responsibilities Support definition of requirements, interfaces, and concept of operations through supporting and/or chairing working group meetings. Understand and communicate prioritization of efforts to multi-disciplinary team. Think abstractly and see the big picture while evaluating technical details Perform technical analyses to develop and validate models of system behavior; identify solutions to complex problems. Serve as a direct interface to both internal and external customers Conduct and support trade s

Recruitment
🔔

Get new model behavior engineer jobs in United States by email

Daily job updates · Unsubscribe anytime