About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en
Jobs in United States
Model Behavior Engineer in United States
2,174 active opportunities · Updated October 2026
Showing
15 jobs
Explore current model behavior engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption
$196K – $230K/yr
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI. This role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Define what “good” looks like (frameworks & rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness,
From $180K/yr
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Senior Staff Rust Developer to join our Platform Convergence Team. This is a hybrid role based in San Jose, CA reporting to the Sr. Director, Software Engineering. Join us to build a new platform from the ground up that can scale hundreds of millions of users with high reliability and low latency. You will design and implement distributed system and core infrastructure components while collaborating closely with various stakeholders. What you’ll do (Role Expectations) Design and build a low-latency, high-throughput data forwarding plane using Rust, leveraging its async/await model for efficient I/O and service-oriented infrastructure Develop distributed, scalable systems with a focus on concurrency, fault tolerance, and messaging Implement and maintain gRPC-based APIs and services to integrate forwarding plane capabilities with control and orchestration layers Optimize system
About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT Multimodal team works across voice, image generation, and other multimodal experiences to turn frontier research capabilities into reliable products. The team connects product usage and failure patterns with research, evaluation, data, inference, capacity, and external partnerships so that model and product improvements translate into better experiences for users. About the Role We are seeking a Technical Program Manager to build the flywheel that helps ChatGPT multimodal products learn from real-world usage and improve quickly. You will lead programs spanning production-signal mining, evaluation and data pipelines, research-to-production parity, multimodal capacity planning, and complex cross-functional dependencies for voice and image-generation launches. You will work closely with product engineering, research, Human Data, inference and capacity teams, safety partners, and external vendors or product partners. Success requires technical depth, strong systems thinking, comfort with ambiguity, and the ability to turn fragmented or manual work into durable mechanisms that teams adopt. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build a system for mining production conversations and product signals to identify representative multimodal workflows, user needs, and failure modes. Establish and maintain evaluations for the highest-priority multimodal behaviors and use cases, with clear coverage, quality standards, and ownership. Package production signals into decision-ready data and
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE The Field Productivity & Enablement Lead is responsible for making our sales motion clear, practical, and repeatable. This person will help define how we sell at Baseten: how managers run the business, how reps qualify and advance deals, how sales works with FDE, product, marketing, and support, and how those expectations get turned into training and day-to-day habits. This is a senior individual contributor role for someone who excels at both strategy and execution. You'll shape the system, but you won't stop at slides or frameworks. You'll build the playbooks, run the training, coach to the behaviors, and help managers ensure the process is followed consistently. We're not looking for someone who wants to force-fit a single methodology onto the business. We're looking for someone who is fluent in MEDDIC/MEDDPICC, Command of the Message, Challenger, and similar approaches, and can take the best ideas from each, adapt them to a technical sales motion, and build an approach that fits how Baseten actually sells. You'll report to the Head of Enablement and Productivity. Frontline managers and reps are your primary customers. RESPONSIBILITIES Manage operating model. Define the core rhythms for frontline managers, including 1:1s, forecast calls, pipeline reviews, deal reviews, and coaching cadences. Create clear expectations for how managers inspect deals, coach reps, and drive consistency across the team. Sale
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. As we push toward AGI, cybersecurity becomes one of the most important and urgent frontiers: the same systems that can accelerate productivity can also accelerate exploitation. As a Researcher for cybersecurity risks, you will help design and implement an end-to-end mitigation stack to reduce severe cyber misuse across OpenAI’s products. This role requires strong technical depth and close cross-functional collaboration to ensure safeguards are enforceable, scalable, and effective. You’ll contribute directly to building protections that remain robust as products, model capabilities, and attacker behaviors evolve. In this role, you will: Design and implement mitigation components for model-enabled cybersecurity misuse—spanning prevention, monitoring, detection, and enforcement—under the guidance of senior technical and risk leadership. Integrate safeguards across product surfaces in partnership with product and engineering teams, helping ensure protections are consistent, low-latency, and scale with usage and new model capabilities. Evaluate technical trade-offs within the cybersecurity risk domain (coverage, latency, model util
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a highly skilled Staff AI Engineer - Multi-Agent Frameworks to join our AI Platform team. In this role, you will play a pivotal part in building a cutting-edge platform that empowers our users to create and deploy sophisticated intelligent agents, with a key focus on enabling collaborative and multi-agentic behaviors . This is a backend-focused role that requires deep expertise in AI, large language models (LLMs), and orchestration software. Key Responsibilities: Design, develop, and maintain a robust platform to enable users to create and manage AI agents and their interactions. Integrate and work with multiple LLMs, ensuring seamless orchestration and scalability for both individual and coordinated agent operations. Leverage orchestration frameworks like LangGraph and others to build complex workflows and pipelines that support diverse agent functionalities, including frameworks for multi-agent coordination . Develop and implement evaluation frameworks for testing AI agents in challenging and complex scenarios, focusing on individual performance and system-level dynamics. Stay at the forefront of AI advancements, incorporating the latest research and technologies into our platform to enhance agent capabilities and collaboration. Collaborate with cross-functional teams, including product managers, designers, and frontend engineers, to deliver a seamless user experience for building and deploying intelligent systems. Address challenging AI privacy scenarios, ensuring compliance with data protection regulations and best practices within agent-based applications. Contribute
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer - Multi-Agent Frameworks to join our AI Platform team. In this role, you will play a pivotal part in building a cutting-edge platform that empowers our users to create and deploy sophisticated intelligent agents, with a key focus on enabling collaborative and multi-agentic behaviors . This is a backend-focused role that requires deep expertise in AI, large language models (LLMs), and orchestration software. Key Responsibilities: Design, develop, and maintain a robust platform to enable users to create and manage AI agents and their interactions. Integrate and work with multiple LLMs, ensuring seamless orchestration and scalability for both individual and coordinated agent operations. Leverage orchestration frameworks like LangGraph and others to build complex workflows and pipelines that support diverse agent functionalities, including frameworks for multi-agent coordination . Develop and implement evaluation frameworks for testing AI agents in challenging and complex scenarios, focusing on individual performance and system-level dynamics. Stay at the forefront of AI advancements, incorporating the latest research and technologies into our platform to enhance agent capabilities and collaboration. Collaborate with cross-functional teams, including product managers, designers, and frontend engineers, to deliver a seamless user experience for building and deploying intelligent systems. Address challenging AI privacy scenarios, ensuring compliance with data protection regulations and best practices within agent-based applications. C
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. As a Formal Verification Engineer at NVIDIA, you will verify the build and implementation of the industry's leading GPUs. In this position, your responsibilities will be to verify the micro-architecture using formal verification tools, define the verification scope, and ensure correctness. You will employ sophisticated formal techniques to acquire sufficiently bounded proofs while working with architects, designers, and pre- & post-silicon verification teams to accomplish your tasks. You will efficiently complete the formal verification effort for the entire project cycle, delivering high-quality results on schedule, and clearly conveying those results to the team. What you will be doing: Identify key behaviors for verification to write clear testplans for sophisticated designs. Implement testplans using the latest formal techniques, including the development of environment assumptions, assertions, and cover properties. Develop abstraction models to overcome complexity challenges and acquire full proofs, or bounded proofs with sufficient coverage. Drive formal tools to realize their best performance. Debug RTL to identify causes of failure scenarios. Contribute to flow and script development
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Job Summary We are seeking a motivated engineer to join the DRAM Systems Engineering team, focusing on the development, evaluation, and optimization of next-generation memory systems for AI accelerators. This role emphasizes research and development across hardware architecture, operating systems, and performance analysis to support Agentic AI inference workloads. If you are ambitious and eager to make an impact in the exciting world of AI and memory systems, this is the perfect opportunity for you! Responsibilities Characterize AI inference workloads and examine memory behavior Build and evaluate tiered memory hierarchies for AI accelerators Study KV cache lifecycles, MoE models, and data placement strategies Compare and optimize explicit versus hardware-assisted data movement Develop, test, debug, and detail system-level and OS components Prototype and evaluate agentic AI systems by building agents and multi-agent workflows using modern frameworks and orchestration patterns (planning, tool use, memory, and context management). Apply these technologies both as workloads under study and as accelerators for internal engineering workflows <h2 style="color:!importan
The Opportunity The world of design is changing rapidly, and the Pro Design team is leading that transformation. We are the Adobe organization behind Illustrator, InDesign, and emerging experiences that connect creativity, collaboration, and AI. Our teams are reimagining what professional design looks like for the next decade - building intelligent, connected tools that empower creators and teams to move faster without sacrificing craft. We are looking for a Senior Business Data Scientist who is creative, analytical, and unafraid to question the status quo and shape the decisions that move key business metrics at scale. Join us and build Adobe’s future products! What you'll Do Map the user funnel and build the metrics, cohorts, and dashboards that Product and Growth rely on to see how users move across free, trial, and paid tiers—and pinpoint where they drop off. Dig into the hard questions (what drives activation, which behaviors predict retention and expansion) and build propensity models for conversion, upgrade, churn, and expansion that feed real-time targeting and in-product nudges. Find and size growth bets and work with Product to ship them. Set north-star, driver, and guardrail metrics with your partners, and stand up multivariate experiments across onboarding, paywalls, in-product prompts, and pricing. What you need to succeed Minimum Requirements: Bachelor's degree in a quantitative field (Statistics, Mathematics, Computer Science, Economics, Engineering, or similar) or equivalent practical experience. 5+ years of experience in data science, product analytics, or a similar quantitative role. Proficiency in SQL and Python (or R) for data manipulation, analysis, and modeling. Hands-on experience designing and analyzing A/B tests and interpre
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 This role is an exciting opportunity to be instrumental in creating magical collaboration workflows within ClickUp. Your focus will be on enabling teams and organizations to work together more effectively, fostering clear communication, coordination, and improved productivity. In this role, you'll have the opportunity to shape the collaboration experience in our products. Responsibilities: Uses AI to gather customer research, generate surveys and tests, and turn insight into clear UX. Prototypes flows, interactions, and logic using Claude Code, Cursor, for rapid iteration with customers and internal stakeholder alignment. Creates high-fidelity UI prototypes that push aesthetics, animation, and visual quality beyond what manual craft alone allows. Automates handoffs to engineering via Figma or vibecoded prototypes built with Claude Code or Cursor. Conduct user research to gain deep insights into user behaviors, pain points, and needs in the collaboration space. Collaborate with cross-functional teams, including product managers, engineers, and researchers, to define product strategy and roadmap for the collaboration experience. Create wireframes, prototypes, and high-fidelity designs that effectively communicate design concepts and interaction models. Conduct usability testing and gather feedback from users to iterate and improve designs. Work closely with front-end developers to ensure the design vision is implemented to the highest standard. Requirements: Minimum of 7 years of experience in product design, with a focus on collaboration or communication products. Strong portfolio demonstrating a track
From $159.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a User Researcher on the Economy team, reporting directly to the Director of Design Strategy & Operations, you will generate actionable qualitative and quantitative insights that shape core economic systems, developer monetization workflows, and immersive brand experiences. Embedded directly on product teams, you will conduct generative and evaluative research centered on how creators build businesses, how brands interact on the platform, and how users participate in virtual economies. You will tackle high-priority features, solve complex behavioral puzzles, and help translate multi-faceted economic infrastructure into intuitive, trustworthy, and scalable experiences. Partnering closely with Product, Product Design, Engineering, Data Science, and Policy, you will ensure user needs directly inform platform decisions while driving mutual value across players, creators, and brands. You will: Execute end-to-end mixed-methods research projects to understand creator monetization, payouts, advertising, and commerce workflows, translating abstract user behaviors into clear product feature recommendations. Conduct foundational research to uncover key user pain points, mental models, and behav
About the Team The Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. Key focus areas include: catastrophic risk, mental health, teen safety and multimodal safety. About the Role Providing access to frontier AI systems raises complex questions around dual-use science and catastrophic risk. How should models respond to requests involving chemical synthesis, biological experimentation, or pathogen research? Where is the boundary between legitimate scientific inquiry and information that could enable misuse? How do we design policies that meaningfully reduce risk without unnecessarily restricting beneficial research? This is a senior role in which you’ll help shape policy creation and development at OpenAI for addressing biological and chemical risks. You will develop structured policy frameworks and taxonomies to guide safe model behavior. This role sits at the intersection of biosecurity expertise, AI safety research, and policy design. You will help ensure that frontier AI systems can support beneficial life sciences research, such as drug discovery, public health, and biosafety, while reducing the risk that these capabilities could be misused. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you’ll: Design and maintain model policies governing chemical and biological risk, defining how models should safely handle dual-use scenarios. Develop structured taxonomies of chemical and biological risk that inform model training data, evaluation benchmarks, and safet
Other cities to consider
More places hiring for this role
Get new model behavior engineer jobs in United States by email
Daily job updates · Unsubscribe anytime