Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Safety AI Systems? As Senior Engineering Manager for Safety AI Systems at Roblox, you'll lead technical efforts and manage a team of experienced engineers to develop innovative AI solutions for multimodal content safety. You’ll oversee machine learning systems, constructing multimodal model architectures, improving data quality, training pipelines, and model performance to address challenges like real-time multi-verse content understanding and advanced moderation with large vision language models, spanning avatars, images, videos, audios, text, code / data models, and their composites. In close collaboration with product, policy, and Trust & Safety teams, you'll design large-scale systems to detect and mitigate abusive behavior before it harms the community. You'll own critical services at massive scale, balancing user freedom with platform civility to protect and empower our users. Your leadership will help ensure Roblox remains a safe, inclusive space for self-expression and shared experiences. You Will Own the vision, technical direction, and execution of machine learning solutions for the Multimodal Safety AI system, ensuring these systems effectively detect and prevent ha
Jobiba hiring network
Model Behavior Engineer Jobs
4,989 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current model behavior engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the role AI Teammates are agents that work like actual team members in Asana — they triage bugs, respond to requests, draft project briefs, conduct research, and handle complex knowledge work across workflows. Unlike chatbots, Teammates are shared team resources that build memory and context across all executions, getting smarter as your team works with them. Currently live with Fortune 500 customers, AI Teammates represent Asana's shift from tracking work to getting work done. We're looking for a Senior Engineering Manager to lead the AI Teammates Platform team — the core engineering team responsible for the execution engine and capability layer that powers AI Teammates. You'll manage a senior team of approximately eight engineers (ICs up to L6) building the systems that determine how AI Teammates reason, execute, and improve over time — from model integration and rollout, to proactive agent behavior, to the developer platform that enables other Asana teams to build on top of AI Teammates. This is a rare opportunity to be at the forefront of Agentic AI. You'll work directly with model partners, lead the core team behind a new applied AI product introduction, and help define how AI agents collaborate with humans in a work management platform used by millions of teams worldwide. This role is based in our San Francisco office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. About Asana AI Asana AI is the company's number one priority. We're building the future of human/AI collaboration — going be
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Safety and Customer Care (SCC) team at Lyft manages over 1.7 million monthly human and AI interactions and serves as Lyft's primary direct touchpoint with riders and drivers. We handle critical infrastructure that powers both human associates and AI agents to make riders and drivers feel safe and comfortable while riding or driving with Lyft, transforming every support interaction into a moment of genuine connection. As a Data Engineer on the SCC team, you will have ownership over the data modeling and pipelines that power SCC’s Associate and AI Agent Platform . Your efforts will be critical to the reliability of our pipelines, execution of third party data integrations, accurate reporting of agents performance, and efficiency improvements that can save millions of dollars / year. You will work cross-functionally to bridge Lyft's business goals with data engineering. Your efforts will allow access to business and user behavior insights, using huge amounts of Lyft data to fuel several teams such as Analytics, Data Science, Engineering, and many others. Responsibilities: Owner of the core data pipeline, responsible for scaling up data processing flow to meet the rapid data growth at Lyft Evolve data model and data schema based on business and engineering needs Implement systems tracking data quality and consistency Develop tools supporting self-service data pipeline management (ETL) SQL and MapReduce job tuning to improve data processing performance Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Collaborate cross-functionally with product, engineering, data science, and marketing teams to understand business problems and align on prioritization and solutions Experience: Bachelor's degree in Compute
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We're hiring a hands-on Engineering Manager to build and lead Replit's Anti-Abuse team from the ground up. This is a foundational 0-to-1 role: you'll define the anti-abuse roadmap, hire a small team of engineers and data analysts, and ship the systems that protect Replit's platform, users, and economics from adversarial actors. You'll partner across Support, Legal, Security, Infrastructure, and the Money and Growth teams to make abuse economically unviable while keeping friction low for legitimate users. Replit sits at the frontier of AI-native abuse. Our platform is a target for phishing and scam hosting, cryptomining, LLM token farming, card and coupon fraud, and increasingly, abuse driven by AI agents themselves. The team you build will define how Replit defends against all of it. What You'll Do Build the anti-abuse roadmap from scratch : Define the threat model, prioritize across abuse vectors (phishing/scam hosting, cryptomining, token farming, payment fraud, AI agent exploitation), and translate it into a shipping plan with clear sequencing and tradeoffs. Design progressive verification and identity infrastructure : Build the "ladder of trust" that gates increasing platform capabilities (referrals, additional credits, access to powerful agent features, Missions) behind escalating verification. This includes a humanity/identity layer that's distinct from user accounts, integrations with KYC-grade verification providers, and the policy engine that decides what level of trust unlocks what behavior. This infrastructure is core not just to promo integrity but to how Replit safely expands agent capabilities over time. Ship as a hands-on EM : Stay in the code. Use the latest AI coding tools (including Rep
#Team Nextdoor Nextdoor (NYSE: NXDR) is the essential neighborhood network. Neighbors, public agencies, and businesses use Nextdoor to connect around local information that matters in more than 350,000 neighborhoods across 11 countries. Nextdoor builds innovative technology to foster local community, share important news, and create neighborhood connections at scale. Download the app and join the neighborhood at nextdoor.com . Meet Your Future Neighbors At Nextdoor, machine learning is one of the most important teams we are growing. Machine learning is starting to transform our product through personalization, driving major impact across different parts of our platform, including newsfeed, notifications, ad relevance, connections, search, and trust. Our machine learning team is lean but hungry to drive even more impact and make Nextdoor the neighborhood hub for local exchange. We believe that ML will be integral to making Nextdoor valuable to our members. We also believe that ML should be ethical and encourage healthy habits and interaction, not addictive behavior. We are looking for great engineers who believe in the power of the local community to empower our members to make their communities great places to live. At Nextdoor, we operate in an AI-first environment and expect every team member to actively use AI tools as part of their workflow. We aren't looking for prompt engineers; we’re looking for people who use tools like Claude, Gemini, ChatGPT, and Glean to challenge their own thinking and take full ownership of AI-assisted outputs. We also offer a warm and inclusive work environment that embraces a hybrid employment model, blending an in office presence and work from home experience for our valued employees. The hiring team will go over these expectations with you if you are being considered for a role near one of our offices in San Francisco, Los Angeles, Chicago, Dallas, New York, and London. The Impact You’ll Make You will be part of a scrappy and
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are looking for an AI Agent Security Architect to function as the primary technical authority for Replit’s autonomous and AI agent security blueprint. In this critical role, you will design, implement, and maintain the runtime defense systems, guardrail frameworks, and sandboxing architectures that govern AI agents executing code, invoking tools, and reasoning across our platform. You will be a key technical contributor—leading high-impact AI security initiatives and bridging the gap between non-deterministic AI behavior and rigorous cybersecurity controls for both engineering and executive leadership. What You'll Do AI Agent Security Strategy & Technical Execution AI Agent Security Blueprint: Define the long-term vision and architectural patterns for securing autonomous agent workflows, Model Context Protocol (MCP) integrations, multi-turn reasoning loops, and multi-agent coordination. Runtime Guardrails & Policy Enforcement: Architect and deploy dynamic input/output guardrail systems, semantic firewalls, and real-time intent verification filters to prevent goal hijacking, system prompt leaks, and indirect prompt injections. Agent Execution & Tool Sandboxing: Partner with Infrastructure and AppSec teams to design secure, short-lived, micro-isolated environments (e.g., microVMs, WebAssembly, container sandboxes) where agents can dynamically execute code, run shell commands, and interact with host operating systems safely. Agentic Threat Modeling & Red Teaming: Conduct specialized threat modeling against non-deterministic systems. Lead automated and manual AI red-teaming initiatives to uncover vulnerabilities in RAG context pipelines, vector stores, and tool-calling interfaces. Identity
About the Team: The OpenAI API team builds the foundation that enables every developer to harness OpenAI’s models safely, reliably, and at scale. We design and operate the systems that power model serving, API access, billing, developer tooling, and enterprise integrations—forming the connective tissue between OpenAI’s research breakthroughs and real-world products. Our mission is to make it effortless for anyone to build with OpenAI technology. We’re responsible for the infrastructure and product layers that allow millions of developers to integrate GPT models, fine-tune behavior, manage data, and deliver transformative experiences to their users. We collaborate across product, research, and engineering teams to ensure that innovation in model capabilities translates directly into value for customers. The API team spans multiple disciplines, including product management, infrastructure engineering, developer experience, and data systems. We care deeply about reliability, scalability, and simplicity—creating tools that let developers focus on their ideas while we handle the complexity of running world-class AI systems. About the Role: We are seeking an experienced Product Manager to define and scale the construction of our data processing, data privacy, billing, and access controls products. You will set strategy and execute on projects like expanding our regional data processing footprint, enabling new inference caching controls in the API or building APIs that make it easier for organizations to manage their spend limits. You will also define the strategy and ship foundational capabilities that ensure customers use OpenAI products securely, privately, and with enterprise-grade controls. This role partners deeply with engineering, security, legal, compliance, finance and leadership to deliver high-trust, enterprise-grade systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to n
About the Team The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to train nuanced safety behaviors, how to make the model robust to bad actors, how to address privacy and security risks, and how to make the model trustworthy in safety-critical situations. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role We’re seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You’ll advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. In this role, you will: Research and implement methods for safety training, reinforcement learning, and adversarial robustness. Develop evaluations, identify model failure modes, and use findings to improve training. Work with research, engineering, security, and policy partners to support safe, reliable deployment. You might thrive in this role if you: Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness. Have a degree in computer science, machine learning, or a related field, and strong deep learning research or engineering skills. Have experience improving model safety for deployment and enjoy collaborative research. Are motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings. Security Requirements Active TS/SCI clearance or equivalent. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefi
About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin
As the Senior Product Manager for the Actions & Automations team, you will own the ecosystem that enables customers, partners, and Datadog teams to build, deploy, and operate AI agents on Datadog. You will drive the strategy and execution for the platform capabilities, developer experience, integrations, and extensibility model that make Datadog the best place to build agents that understand and act on production systems. Modern engineering organizations are entering a new era where software is not only monitored and operated by humans, but increasingly by AI-powered agents. As agentic workflows reshape how teams build, operate, secure, and troubleshoot systems, customers need a platform for creating specialized agents, connecting them to business and engineering systems, governing their behavior, and extending them to solve unique organizational problems. You will define and build the ecosystem that makes this possible. Agent Builder sits at the intersection of Datadog's products, AI capabilities, and ecosystem strategy. You will have the opportunity to work across the breadth of the Datadog platform, partner with teams throughout the company, and help establish Datadog as the foundation for operational AI. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define the vision, strategy, and roadmap for Datadog's Agent Builder platform and ecosystem. Own the core platform capabilities that enable customers and partners to create, customize, deploy, and manage AI agents. Drive the extensibility model for agents, including integrations, tools, actions, context sources, APIs, SDKs, and developer workflows. Shape how agents perform actions across Datadog products and third-party systems. Partner closely with AI, platform, infrastructure, and product tea
About the Team: The OpenAI API team builds the foundation that enables every developer to harness OpenAI’s models safely, reliably, and at scale. Our mission is to make it effortless for any developer to build transformative products with OpenAI’s models. We’re responsible for the infrastructure and product layers that allow millions of developers to integrate our models, fine-tune behavior, manage data, and deliver experiences to their users. We collaborate deeply across product, research, and engineering teams to drive innovation at the model layer and then directly translate that into value for customers. About the Role: As the Agents Product Manager for the API team, you'll be at the forefront of defining and guiding the future of how developers build agentic applications on top of our AI models. You’ll set clear priorities and drive impactful improvements to model capabilities, balancing user needs, safety considerations, and technical innovation. This role is perfect for a proactive, technically adept PM who thrives on solving challenging, ambiguous problems through structured product thinking and close collaboration with customers, engineers, and researchers. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Deeply understand problems faced by agent builders and identify opportunities where our products and models can make building agents faster, more intuitive, more reliable, and more powerful. Define strategic priorities and roadmap for improving agentic infrastructure for API users, focusing on user outcomes and emerging capabilities. Partner with research and engineering teams at a technical level to translate those priorities into developer products and features (SDKs, APIs, and more). Deliver quickly while maintaining a high bar for product quality and user experience. You might thrive in this role if you: Have 5+ years of product management or related industry experience. Proven track record of b
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About the Team At Embedded Insights, we find the best machine learning opportunities for external products and internal systems, and collaborate with cross-functional partners to bring them to life. We are a central team of Machine Learning Engineers and Data Scientists. We embed with partner teams to build and apply machine learning models that improve internal decision-making and power the Plaid product suite. About the Role You will be the first Data Scientist on the Embedded Insights team, part of Plaid’s Data organization. You will establish the analytics and metrics backbone for a team supporting a diverse set of internal and external products. You will help drive better decision-making, support machine learning model development, and contribute directly to the health of the Plaid network and the quality of Plaid’s products. Your day-to-day work will include: Analyzing entities across the Plaid network to understand behavior and identify opportunities, anomalies, and risks. Creating foundational metrics, dashboards, and monitoring systems that provide a clear view of network health and machine learning model performance. Evaluating the value and performance of machine learni
About the Team OpenAI’s Education team is building products that advance how people learn with AI. The team works across higher education institutions, K-12 districts, and country-level partnerships, including applied research on how AI affects learning and cognitive outcomes. The team owns owns ChatGPT Edu, ChatGPT for Teachers, and related product/research work. The team partners closely with go-to-market, research, Consumer Learning, and model teams to turn education-specific insights into product experiences that can improve ChatGPT more broadly. Some of our recent work: New Education Plugins for ChatGPT Work and Codex New tools for understanding AI and learning outcomes Education for countries Advancements in higher education Early product work - Introducing Study Mode About the Role We’re looking for a hands-on Tech Lead Manager to lead and manage a team of senior full-stack engineers building AI-native learning experiences in ChatGPT. This person will combine technical execution, product judgment, and people leadership: they will write and ship code, manage engineers, and help shape the product direction for how students and Educators use AI. In This Role, You Will Lead and manage a team of three senior full-stack engineers. Build product experiences for ChatGPT Education, ChatGPT for Teachers, and AI-native learning workflows. Partner with research teams on field studies, randomized control trials, classifiers, data pipelines, and cognitive-outcome measurement. Collaborate with Consumer Learning and model teams to translate education insights into broader ChatGPT behavior and product improvements. Drive execution across product, engineering, research, go-to-market, and partner teams. Help define product strategy, priorities, and delivery plans for a new product pod. You Might Thrive In This Role If You Have several years of direct people-management experience with engineers. Are still highly technical and comfortable doing IC engineering work. Have strong pr
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the role In this role, you will collaborate closely with researchers and engineers on the Learned Behavior teams to tackle plan generation challenges in autonomous driving. You’ll apply state-of-the-art generative modeling techniques—ranging from cutting-edge diffusion, flow matching, energy-based models, and SoTA algorithms—in order to develop novel solutions that generate safe, comfortable, and efficient driving behaviors in the most challenging real world situations. Beyond core research, you’ll own the end-to-end lifecycle of your models, productizing them for robust, real-world autonomous driving deployments on a global scale. About the Work Develop and scale state-of-the-art generative models—especially diffusion architectures, flow-matching techniques, and energy-based models —for autonomous plan generation. Build generative models with foundation models. Leverage large language models and world foundation models for r
The Opportunity MongoDB’s partner ecosystem — systems integrators, ISVs, and the major cloud providers — is a strategic growth engine for the business. We’re looking for a Senior Manager, Sales Plays and Offerings to design the joint go-to-market motions that turn partner relationships into pipeline and revenue. This is a highly cross-functional, strategic role: you’ll build the sales plays and joint offerings themselves, partner with Enablement to get the field and our partners ready to sell them, instrument how they perform, and work with the Programs lead to make sure incentives reward the behavior we want to see. You will report directly to the VP of Partner Strategic Operations and act as a connective layer between Partnerships, Sales, Enablement, and Programs — translating ecosystem strategy into repeatable, measurable, field-ready motions. We are looking to speak to candidates who are based anywhere in the US for our hybrid working model. What You'll Do Build sales plays and joint offerings Design and package partner sales plays and joint solution offerings with priority ISVs, SIs, and cloud partners — defining the joint value proposition, target segment, competitive positioning, and playbook for how field and partner sellers execute it Partner with Product Marketing, Solutions Engineering, and partner counterparts to validate technical integration stories and translate them into a compelling, sellable narrative Prioritize which plays to build and scale based on market opportunity, partner readiness, and alignment to MongoDB’s strategic pillars (e.g., AI, migrations, industry verticals) Own the lifecycle of each play from concept through launch, iteration, and eventual retirement or refresh Drive partner and field enablement Partner closely with the Enablement team to translate each sales play into field- and partner-facing assets: pitch decks, battlecards, demo scripts, certification conte
Get new model behavior engineer jobs by email
Daily job updates · Unsubscribe anytime