Jobs in United States

Model Behavior Engineer in United States

2,174 active opportunities · Updated October 2026

Explore current model behavior engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety. Key focus areas include: catastrophic risk, mental health, teen safety and multimodal safety. About the Role Providing access to frontier AI systems raises complex questions around dual-use science and catastrophic risk. How should models respond to requests involving chemical synthesis, biological experimentation, or pathogen research? Where is the boundary between legitimate scientific inquiry and information that could enable misuse? How do we design policies that meaningfully reduce risk without unnecessarily restricting beneficial research? This is a senior role in which you’ll help shape policy creation and development at OpenAI for addressing biological and chemical risks. You will develop structured policy frameworks and taxonomies to guide safe model behavior. This role sits at the intersection of biosecurity expertise, AI safety research, and policy design. You will help ensure that frontier AI systems can support beneficial life sciences research, such as drug discovery, public health, and biosafety, while reducing the risk that these capabilities could be misused. Our relevant publications: Preparedness framework Preparing for future AI capabilities in biology Safety evaluations hub OpenAI GPT5 System Card Evaluating Fairness in ChatGPT Improving Model Safety Behavior with Rule-Based Rewards OpenAI Model Spec Your Responsibilities: Design and maintain model policies governing chemical and biological risk, defining how models should safely handle dual-use scenarios. Develop structured taxonomi

AWSGitRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Core Models team shapes how our models interact with people. We view the model as the product itself, aiming for intuitive experiences that exceed user expectations and feel like magic. About the Role As a Model Designer, you’ll have an outsized impact on how our models interact and resonate with users. You’ll strike a delicate balance between maximizing the model’s capabilities, reading in between the lines in user queries to understand how best to help, and upholding user trust. We’re looking for people who are passionate about the intersection of design, technology, and user experience — and are up for the challenge of defining new human-AI interaction paradigms. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you’ll join a small team in evolving and expanding the model design function. You will: Collaborate closely with researchers to understand, predict, and design model behavior. Partner with product managers and designers across the company to ensure a cohesive voice and approach. Proactively identify ways to improve our models based on product sense, user feedback, quantitative insights, and the research roadmap. Come up with creative strategies for collecting high-quality data. Do whatever needs to be done to make our models better. You might thrive in this role if you: Possess exceptional taste, creativity, and writing skills, allowing you to craft responses that delight users. Don’t mind ambiguity — you’re happy to throw yourself into a new, unfamiliar environment, build relationships, define a problem, and make progress. Love experimentation, and are willing to test and reject new ideas when the results don’t pan out. Exhibit high levels of empathy and self-awareness required to serve everyone in the world. Enjoy tackling profound and often philosophical questions while always driving towards clarity. Demonstrate technic

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI is at the center of some of the highest-impact multimodal work in AI. ChatGPT serves a massive global audience, and enables diverse interactions via text, speech, and visuals. As interactive surfaces grow, models also need to adapt to emerging harm, understand user intent and situational context, and respond appropriately. The Chat and Multimodal Safety team is responsible for ensuring that OpenAI’s increasingly multimodal models and products behave safely across these experiences. We develop the research, training methods, and evaluations needed to make these experiences safe. Our work sits at the frontier of responsibly deploying powerful AI, in close partnership with Personal AGI, io, model training, and product teams. About the Role As a Researcher on the Chat and Multimodal Safety team, you will help shape how frontier models perceive and reason the world, and translate that understanding into safe behavior. We’re looking for people who combine deep technical expertise with strong safety judgment. Strong candidates often bridge perception and language: they may have built vision-language models, worked on modality fusion or image encoders, developed multimodal post-training or evaluations, or advanced safety for image, video, or audio systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Define and advance multimodal safety research for text, vision, and audio, connecting perception and semantic understanding to safe model behavior. Build training and evaluation methods for VLMs, including post-training, safety evals, and interventions that help models respond safely and appropriately in varied contexts. Collaborate closely with Personal AGI, Consumer Devices, and product/model teams to translate research into safer ambient, embedded, and personalized multimodal experiences. You might thrive in this role if you:

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability , which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor). Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research.We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable. About the Role We’re looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work. As a researcher on the Alignment team, you will design and run experiments that improve our understanding of model monitorability. You will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make those questions measurable, and help translate findings into practical oversight and training recommendations. You may also help develop new monitoring models or methods and apply them to OpenAI’s largest training runs. This role is especially well

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. Within Safety Systems, the Model Policy team works to ensure that increasingly capable models behave safely and reliably in real-world environments. We investigate emerging model failures, define the behavior models should exhibit instead, and develop the data, evaluations, monitoring, and safeguards needed to improve and validate that behavior. Our work connects alignment research with the practical challenges of training and deploying frontier models. About the Role In this role, you will shape how OpenAI understands and addresses real-world risks that emerge from model misalignment as models become more autonomous and operate over longer horizons. You will investigate how misaligned behavior emerges across extended trajectories - including when models persist toward the wrong objective, take unsafe shortcuts, lose track of instructions, exploit weaknesses in their environment, or circumvent constraints - and translate these insights into behavioral policies, evaluations, monitoring, and safeguards. This role is ideal for someone who wants to turn alignment and safety concerns into concrete, empirically grounded improvements to frontier AI systems. Your Responsibilities: Identify vulnerabilities that emerge as models interact with tools, data, and external systems, and translate them into model- and system-level safeguards. Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned behavior. Build frameworks for understanding harmful outcomes arising from model misalignment. Identify the underlying behaviors and system conditions that drive those outcomes. Turn findings into policy frameworks, evaluation criteria, online measurement and safeguards. Develop human data campaigns and gold sets to ground measurement and evaluation of eme

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The Safety Systems team is responsible for various safety work to ensure our best models can be safely deployed to the real world to benefit the society and is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency. The Model Safety Research team aims to fundamentally advance our capabilities for precisely implementing robust, safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to enforce nuanced safety policies without trading off helpfulness and capabilities, how to make the model robust to adversaries, how to address privacy and security risks, and how to make the model trustworthy in safety-critical domains. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role OpenAI is seeking a senior researcher with passion for AI safety and experience in safety research. Your role will set directions for research to enable and empower safe AGI and work on research projects to make our AI systems safer, more aligned and more robust to adversarial or malicious use cases. You will play a critical role in shaping how a safe AI system should look like in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI. In this role, you will: Conduct state-of-the-art research on AI safety topics such as RLHF, adversarial training, robustness, and more. Implement new methods in OpenAI’s core model training and launch safety improvements in OpenAI’s products. Set the research directions and strategies to make our AI systems safer, more aligned and more robust. Coordinate and collaborate with cross-functional team

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About The Team The Data Understanding team is responsible for creating the high quality datasets and their quantized representation for OpenAI. This includes synthesizing data, building VQ representations, and processing, filtering, deduplication, quality control, and tokenization so it can be used effectively in big model training runs. About The Role We're looking to advance how OpenAI builds and understands pretraining data at scale. You'll treat data quality and curation as core research problems: developing new methods to select, combine, and transform data; creating datasets that improve model capabilities; and designing rigorous experiments to understand how data choices and interventions affect model learning and downstream behavior. You'll work closely with frontier models and web-scale data to build evidence for which approaches work and why, then translate successful research into scalable data processing pipelines We Expect You To Have a strong track record of new or improved ML ideas, through publications, projects, or applied research. Own and drive a research agenda, from choosing the right problems to carrying long-running work through to impact. Be excited by OpenAI’s empirical, collaborative approach to research. Nice To Have Thoughtfulness about AI’s impact, including privacy, provenance, and data quality. Experience building high-performance deep learning or large-scale data processing systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the team The Intelligence and Investigations team seeks to rapidly identify and mitigate abuse and strategic risks to ensure a safe online ecosystem in close collaboration with our internal and external partners. Our efforts contribute to OpenAI's overarching goal of developing AI that benefits humanity. This role focuses specifically on AI Safety: understanding and mitigating risks created or amplified by increasingly capable AI systems. It is not a cybersecurity, information security, or corporate security role. The Strategic Intelligence & Analysis (SIA) team provides safety intelligence for OpenAI’s products by monitoring, analyzing, and forecasting real-world abuse, geopolitical risks, and strategic threats. Our work informs AI safety mitigations, product decisions, and partnerships, ensuring OpenAI’s tools are deployed responsibly across critical sectors. About the role We are looking for a Frontier AI Risks Lead to help us understand potential harms and misuse of AI in a time of rapid, sustained change. We seek to understand how developments in AI could intersect with misuse and abuse, accelerating existing harm areas and creating novel risks. We seek to scan available signals and use strategic foresight methodologies to enable proactive detection and mitigation of frontier AI risks. This is an AI safety role focused on frontier and systemic risks, including model misalignment, recursive self-improvement (RSI), multi-agent interaction, loss of control, runaway agents, and related emerging failure modes. In this role, you will help provide a strategic-level perspective on a range of frontier AI safety areas, producing actionable understanding of issues relevant to OpenAI’s platforms, systems, and broader mission. Utilizing mixed quantitative and qualitative methodologies, you will spot early warning signs, pull threads on potentially concerning behavior, and turn weak signals into clear, prioritized risk calls. You will focus on upstream ecosystem sc

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions optimized for advanced AI workloads. We collaborate across research, software, and external hardware partners to design and deploy next-generation AI systems at scale. Our team works closely with silicon vendors and system partners to evaluate emerging technologies, validate performance characteristics, and ensure that hardware capabilities translate effectively to real-world AI workloads. About the Role We are seeking a 3P Hardware Architecture Expert with deep expertise in GPU and accelerator architectures to engage directly with silicon vendors and guide hardware decisions for AI infrastructure. In this role, you will evaluate architectural tradeoffs across compute, memory, and interconnect systems, translating vendor specifications into real-world workload impact. You will play a critical role in early silicon evaluation, benchmarking, and performance validation, helping ensure that next-generation hardware meets the needs of our workloads. This role is highly hands-on and requires both deep technical understanding and the ability to engage at a high level with partners such as NVIDIA and AMD on architectural direction and design tradeoffs. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Engage deeply with silicon vendors (e.g NVIDIA & AMD) on GPU and accelerator architecture tradeoffs. Analyze and interpret performance, power, and efficiency characteristics of next-generation hardware. Translate vendor specifications into expected real-world performance for AI workloads. Evaluate architectural aspects including: compute throughput and utilization memory systems (HBM, cache hierarchies, bandwidth constraints) data types and precision tradeoffs (FP16, BF16, FP8, etc.) interconnect and scaling behavior. Run benchmarks and profiling to validate hardware performance a

AWSRestAIRust
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. About the Team The Finance Data Science team owns the forecasting systems that power Snowflake’s revenue planning and long-term financial strategy. Our work supports corporate planning, executive decision-making, and investor reporting, and we partner closely with Product and Sales to understand customer behavior and product impact. We operate at the intersection of machine learning, statistical research, and corporate finance, building production-grade forecasting infrastructure that is foundational to how the company plans and operates. The Role As a Senior Data Scientist, you will independently lead high-impact modeling initiatives and build production-ready forecasting systems for core financial metrics. You will work on complex, open-ended problems at the intersection of machine learning and business strategy, translating real-world financial questions into rigorous, scalable models. What You’ll Do Design and implement advanced time-series and probabilistic models (e.g., hierarchical models, state-space models, Bayesian approaches, multivariate forecasting). Contribute to internal tooling and shared infrastructure that enables scalable forecasting and analytics. Establish best practices for model evaluation, backtesting, uncertainty quantification, and scenario simulat

PythonSQLMachine LearningAI
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. At Tenstorrent, we build open, state of the art compute for real workloads and real developers. You will own CPU focused test generator development and verification strategy, shaping how our out-of-order RISC-V CPUs are validated against complex ISA and microarchitectural behavior. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You bring 8+ years in CPU design verification, test generation, or closely related CPU validation work. You have led development of test generators for x86, ARM, or RISC-V ISA environments. You understand CPU ISA behavior, privileged architecture, and high-performance out-of-order CPU microarchitecture. You are comfortable building tools, stimulus, and automation that scale verification across large CPU programs. You communicate clearly across design, DV, architecture, emulation, and post-silicon teams. What We Need Lead development of CPU core-level test generators for high-performance out-of-order RISC-V cores. Own generator strategy, infrastructure, and methodology for ISA and microarchitectural verification in both pre-silico

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA builds the silicon behind AI, accelerated computing, and graphics. Every watt of performance and every degree of thermal headroom traces back to decisions made in power, performance, and thermal architecture. We are the Silicon Co-Design Group (SCG). We identify, own, and drive system-level co-design ideas. We start with initial concepts and advance to product differentiation across NVIDIA's roadmap. We are hiring a Principal System Power Management and Performance Architect who operates at the ambiguous boundary where workload behavior, silicon capabilities, firmware policies, and platform constraints collide, and who turns that ambiguity into architecture that survives across multiple silicon generations. SCG scope spans architecture, design, software, operations, platforms, and productization. This role shapes system, platform, and data center features and behavior, and partners with teams across NVIDIA. What You'll Be Doing: The work here is rarely well-defined when it arrives. You will be given problems that appear to be performance gaps or power anomalies and encouraged to build a framework for solving them, not just tackle a single instance. Define the multi-generation roadmap for system-level power and performance features, grounded in prototyping, use-case analysis, and cost/benefit trade-offs across segments. You will decide what problems are worth solving and why. Own the architecture and integration strategy for HSIO power management, DVFS, P-states, and low-power features. Your decisions improve product performance, power, and reliability across product lines — not just the current program. Lead system-level boot and IST architecture defining how power and clock domains initialize, sequence, and recover across complex multi-IP systems where the interaction space is large and the failure modes matter. Drive power management strategy at data

D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Senior Product Manager - Search Datadog’s Search team helps users – both human and agents – find answers to their questions. Search is a critical function in Datadog, and touches many product surfaces: the query editors in product homepages and dashboards, search bars for finding relevant assets, the global cmd+k navigation search, and the MCP tools that our Bits AI agent uses to respond to natural language prompts. Search is a full-stack team: owning user-facing search components, backend search ranking systems, and machine learning models to produce recommendations. The Search team is relatively new, and still growing. We have recently built out agentic search tools, and we are looking for a leader to help us expand to ambitious orchestration systems that deliver accurate results, and proactive recommendations across both keyword and semantic search. Beyond this, you’ll have room to influence how Datadog thinks about Search as a strategic surface. What you’ll do: Define and deliver how Datadog's AI agents discover the right context and tools to answer natural-language questions accurately and at scale Stay on top of industry trends in UI and agentic search experiences and capabilities Collaborate with Applied AI teams to integrate ranking, personalization, and recommendation models that scale across both human and agentic users Define and monitor KPIs for search quality, adoption, and downstream impact on user productivity; use them to drive data-informed decision making Engage directly with customers and internal product teams to deeply understand search journeys across query editors, global navigation, and natural-language agent interfaces Who you are: You have experience with search, ranking, or recommendation systems You are familiar with or very interested in MCP servers and differences between human and agentic UX You have a sharp eye for design and strong opinions on the micro-interactions — keyboard navigation, autocomplete behavior, loading

RestMachine LearningAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Like every team at OpenAI, the Marketing team contributes to our broader mission of ensuring responsible and widespread adoption of artificial intelligence. With that aim in mind, we are responsible for developing and executing strategies that drive awareness, engagement, and usage for OpenAI’s products and platform amongst our core audiences. Our focus extends beyond just promoting product features; we aim to provide valuable insights and resources that help our users make the most out of AI technologies. The Marketing Insights & Analytics team aims to understand the markets, audiences, behaviors, attitudes and needs to shape outreach strategies and inform product and business decisions. We’re the strategic partner that discovers insights, specifically: foundational understanding, marketing strategy, thought leadership, measurement and tracking. About the Role You will be the Market Research Lead that translates insights into actionable marketing plans and executions, drive business results, and orient the team towards strategic thinking to drive brand outcomes. This is critical as we hope to reach millions of users and businesses worldwide. This role can be based in San Francisco, CA or NYC, NY utilizing a hybrid work model (3 days per week in-office). In this role, you will: Play an integral part of Marketing and be a strategic partner to brand, product marketing, creative, product and beyond. Design, execute, and deliver high-impact research using mixed methods. Work on a diverse portfolio of business challenges - how do we reach the next billion people, how do we grow internationally, how do we position our suite of products, and how do we ship joy to our customers. Translate business questions into research plans. Translate data into clear, compelling narratives. Develop deep empathy for developers, builders, and technical decision-makers, understanding their workflows, motivations, and pain points across the lifecycle Shape go-to-market str

AWSRestAIGo
C
📍 Field Louisiana, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. The District Leader, Rx plays a critical role in cultivating a culture of clinical and business excellence in retail pharmacies across their market. With clinical and business oversight of approximately 20 retail pharmacies and an average team headcount of 200 reports, the District Leader, Rx has ultimate responsibility for patient safety and business success in their market. An inspiring coach and leader of people, the District Leader, Rx is focused on building and developing a team of Pharmacy Managers who grow the business through their consistent delivery of unparalleled care and patient connection. Specifically, the District Leader, Rx partners with and coaches their team to deliver on the core business with excellence, embrace change in an evolving healthcare ecosystem, and launch new products, programs, and services to keep our communities healthy. A licensed Pharmacist themselves, the District Leader, Rx brings deep knowledge of pharmacy workflow and clinical programs to help their teams identify and address performance opportunities, grow the top-line through the delivery of clinical services, and abide by all legal and regulatory guidelines with safety of our colleagues and patients top of mind. A model for all CVS Retail Pharmacists and Technicians, the District Leader, Rx also lives our Heart At Work behaviors, and sets the bar for Pharmacy Managers by their examples. To enable sustainable market su

ExcelRecruitmentPayrollCustomer Service
🔔

Get new model behavior engineer jobs in United States by email

Daily job updates · Unsubscribe anytime