Jobiba hiring network

Model Behavior Engineer Jobs

4,989 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current model behavior engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Alignment team at OpenAI is dedicated to ensuring that our AI systems are safe, trustworthy, and consistently aligned with human values, even as they scale in complexity and capability. Our work is at the cutting edge of AI research, focusing on developing methodologies that enable AI to robustly follow human intent across a wide range of scenarios, including those that are adversarial or high-stakes. We concentrate on the most pressing challenges, ensuring our work addresses areas where AI could have the most significant consequences. By focusing on risks that we can quantify and where our efforts can make a tangible difference, we aim to ensure that our models are ready for the complex, real-world environments in which they will be deployed. The two pillars of our approach are: (1) harnessing improved capabilities into alignment, making sure that our alignment techniques improve, rather than break, as capabilities grow, and (2) centering humans by developing mechanisms and interfaces that enable humans to both express their intent and to effectively supervise and control AIs, even in highly complex situations. About the Role As a Research Engineer / Research Scientist on the Alignment team, you will be at the forefront of ensuring that our AI systems consistently follow human intent, even in complex and unpredictable scenarios. Your role will involve designing and implementing scalable solutions that ensure the alignment of AI as their capabilities grow and that integrate human oversight into AI decision-making. This role is especially well suited for someone who can move from an ambiguous model-behavior question to a concrete experimental setup: formulate the hypothesis, build the evaluation or intervention, run the experiment, analyze the result, and decide what the evidence supports. This role may be based in San Francisco or London, subject to team needs and location approval. In this role, you will: We are seeking research engineers and res

typescriptpythonaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Artifacts, you will train frontier models to create polished, useful work products: documents, spreadsheets, slide decks, dashboards, reports, analyses, and other interactive or editable artifacts. You will help teach our models to move from a vague user goal to a finished artifact with strong structure, visual taste, domain judgment, correctness, and low latency. This work will require owning improvements across our post-training stack, including RL, data pipelines, graders, reward signals, evals, and behavioral analysis. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you will: Design and run experiments that improve agentic model behavior for complex so

awsrestmachine learning
View job →

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a

awsrestai
View job →

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Intelligence and Investigations team is dedicated to ensuring the safe, responsible deployment of AI by rapidly detecting and mitigating abuse. Our team leverages the latest testing methodologies to uncover vulnerabilities and emerging threats, helping safeguard OpenAI’s products and users. We work closely with cross-functional partners across product, policy, and engineering to drive a comprehensive defense strategy against evolving adversarial challenges. About the Role As a Red Team Specialist focused on cyber, you will help answer two practical questions: What cyber capabilities can our models provide to real-world attackers, and do our safeguards remain effective when those attackers use increasingly sophisticated techniques? The role combines scaled evaluation with expert-driven testing. You may bring deeper experience in cybersecurity and use that expertise to judge whether a model’s behavior meaningfully changes attacker capability. Alternatively, you may bring deeper experience in model evaluations, automation, or agentic harnesses and apply those skills to building rigorous cyber testing. We do not expect every candidate to be equally deep in both areas, but successful candidates will have a strong foundation in one and enough fluency in the other to work effectively across the boundary. Most of your work will focus on model cyber capabilities and safeguards; you will also spend a portion of your time testing novel abuse risks in agentic systems. This role is located in San Francisco, CA or Seattle, WA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and run rigorous evaluations of model cyber capabilities and safeguards, including policy adherence, correct refusal, over refusal, and resilience to jailbreaking and other adversarial techniques. Conduct hands-on testing to understand what models can enable when used by experienced security practiti

awsrestai
View job →
B
11 days ago

Lead Systems Integrator Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Lead Systems Integrator (Level 4) to support an Air Dominance Fixed Wing Proprietary Program in St. Louis, MO . Step into a fast-paced, cutting-edge program where your expertise will drive the design, integration, and testing of advanced, cloud-based systems that empower mission planning, debrief, and tactical Command & Control (C2) solutions for fixed-wing platforms. In this role, you will lead the development, analysis, and integration of innovative engineering solutions for critical Ground System capabilities and ensure seamless support for ground systems from initial design through to final delivery. You will partner closely with the team’s test lead as well as the overall chief integrator who oversees all activities related to design, integration, and test. You will work with a high-performing, cross-functional team in an agile environment, driving next-generation capabilities from design through delivery. This role offers the chance to innovate with open-architecture, model-based designs and collaborate across disciplines to support critical defense missions. You champion best practices, processes, and standards for the integration team, mentor systems and electrical engineers, and coach solution leads on execution. If you’re passionate about advancing mission-critical systems and thrive in a collaborative, fast-moving environment, this is your opportunity to make a significant impact. Position Responsibilities Lead team of systems and electrical engineers in behavior modeling and requirements decomposition for system, sub-system, and software design

recruitment
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We are looking for a Senior/Staff Software Engineer to serve as a technical leader for Nuro’s ML Data engine. You will sit at the critical intersection of Autonomy, Machine Learning, and Infrastructure, acting as an architect for the systems that feed our autonomy AI models. In this role you will be a member of the Autonomy team responsible for executing the technical strategy for transforming massive amounts of autonomy data into high-value training signals for autonomy decision making. You will design and build data products for autonomy researchers, develop queries for rare "needle-in-a-haystack" scenarios, and trigger labeling and data ingestion workflows without human intervention. You will partner directly with Autonomy ML researchers to understand their data needs, collaborate with infrastructure teams to define the right data interfaces and APIs, and build robust data selection, simulation, and introspe

pythonmachine learningai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The behavior team at Nuro develops the Nuro Driver’s prediction and planning systems to enable safe, driverless autonomy. We are looking for strong software engineers to research, develop, and implement technologies for Nuro’s planning stack to empower all rides on all roads. This encompasses building a generalizable and scalable ML planner that can power L4 driving for robotaxi applications as well as serve as an L2 solution for the largest auto manufacturers in the world. Operating domains of the Nuro Driver™ vary from structured surface streets and highways to more unstructured areas such as parking lot driving, parking and multi-point turns in busy traffic situations. You’ll be developing state-of-the-art algorithms to enable the Nuro Driver to safely and reliably plan and navigate complex situations in a human-like manner. You will work on intelligent strategies on how to leverage diverse data to have the biggest impa

pythonmachine learningai
View job →
N
Nuro
📍 Mountain View• Full-time• From $176.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role As a Senior/Staff Software Engineer working on driving behavior verification, you are responsible for implementing metrics that evaluate the end-to-end behavior of the Nuro Driver. These metrics will be used to quantify the safety of the driving behavior in our target ODD. This requires prior experience with the development or verification of behavior planning/prediction systems for robots, and a collaborative nature to work closely with a variety of teams across Nuro: Systems, Onboard Software, Simulation, Product, and Operations. About the Work Develop and implement in Python generalizable metrics to verify the driving behavior of an autonomous vehicle. Leverage a combination of machine learning (ML) models and safety metrics from literature to evaluate the end-to-end driving behavior. Evaluate these metrics on a variety of tests: synthetic and log simulation, on-road logs, closed-course testing data, and third-party acc

pythonmachine learningai
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Systems Engineering team is responsible for the requirements, architecture, and validation of autonomous driving capabilities across engineering disciplines. This includes designing performance metrics, evaluation methods, and criteria for success, which the team then drives cross-functional via requirement definition and system validation. Systems Engineering works at the intersection of hardware, software, and robot operations, with a deep understanding of technologies in all three. We are a small, high-impact team that sets the checkpoints for autonomy deployment. About the Role The Senior Systems Test Engineer (STE), Autonomy Behavior role is responsible for transforming Nuro's Verification & Validation (V&V) ecosystem. You will focus on the technical implementation, standardization, and automation of V&V processes. This involves creating a unified, automated architecture for validation , designing and

pythonaic++
View job →
N
Nuro
📍 Mountain View• Full-time• From $235K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We are looking for a Senior/Staff Machine Learning Engineer to be a technical leader on Nuro’s Behavior & Planning team. Our team owns how the Nuro Driver behaves on the road: prediction, decision making, and planning, and is responsible for turning Nuro’s large-scale driving data into safe, comfortable, and natural driving behavior. In this role you will bring strong, general machine learning expertise to some of the hardest problems in autonomy and drive them from research through to deployment on real vehicles. You’ll work at the frontier of applied ML spanning areas such as foundation and world models, LLM/VLM reasoning, reinforcement and imitation learning, generative and diffusion models, and transformer-based prediction and planning. You’ll apply this knowledge and experience to make our driving generalize as we scale across new geographies as we expand throughout the U.S. and globally, and across new vehi

pythonrestmachine learning
View job →
E
12 days ago

We build and operate the compute infrastructure our researchers run on, supporting large-scale processing of historical market data and model training on our own hardware across multiple data centers. Our environment includes bare-metal Linux, virtualization, storage, and GPU clusters, where performance, reliability, and predictable system behavior are critical. Our Infrastructure team covers monitoring and automation, distributed storage, hardware and OS provisioning, GPU clusters and workload scheduling, high-speed networking, L2/L3 Linux support, and security engineering. Engineers here own their tasks end to end, so there's room to go deeper in your area and pick up the parts you haven't touched yet. We’re looking for a Linux Infrastructure Engineer who can work hands-on with server and cluster environments, from deployment and configuration to performance tuning, troubleshooting, and ongoing improvement What You’ll Be Doing: Deploying, configuring, and maintaining Linux-based bare-metal servers across our data centers Building and operating clustered environments, including virtualization, storage, GPU compute, and database clusters Troubleshooting complex Linux, hardware, networking, and cluster-level issues Performance tuning for throughput, latency, stability, and resource utilization Monitoring infrastructure health and performance, identifying bottlenecks, and preventing recurring issues Supporting the full server lifecycle: provisioning, setup, upgrades, and maintenance Improving reliability and predictability during failures, maintenance, and scaling Automating provisioning, configuration, and operational tasks, primarily using Ansible and scripting What We Look For In You: Strong hands-on Linux administration and troubleshooting experience Production experience with on-premise, bare-metal infrastructure Good understanding of Linux performance and bottleneck analysis Experience with: infrastructure monitoring and troubleshooting production issues,

REMOTElinuxansible
View job →
D
1mo ago

The Behavior AI team builds the AI-based anomaly detection behind Datadog's security products. Our models learn what normal looks like across the billions of logs, events, and telemetry records flowing through the platform every second, and they flag the behavior that does not fit, on every record, in real time, at a cost that makes sense at our scale. What we build does not ship to a single feature. The same models power detection across many of Datadog's security products at once, so the work has impact well beyond any one team. Large general-purpose models are too slow and too expensive to run in that path, so we take the opposite approach: small, custom models, designed for high-throughput stream processing and optimized to run cheaply on every record. We are hiring a Senior Applied Scientist to build these models from start to finish. You will contribute to designing the architecture, training at scale, and the optimization work that takes a model from training to running efficiently on production traffic. This optimization requires a deep understanding of the constraints imposed by both the software and the hardware, together with the applied mathematics to work within them: often it comes down to finding a mathematical reformulation that fits those constraints better, and that judgment can decide whether a model reaches production at all. The work involves a number of open questions. How do you obtain most of the quality of a large model from one that is far smaller and cheap enough to run on the full stream? Where is it worth trading exactness for speed, and how do you reason about the error you accept? How do you make a small model's outputs clear enough that the detection engineers and analysts who rely on it can trust what it reports? If these are the problems you want to work on, we would like to hear from you. At Datadog, we place value in our office culture: the relationships and collaboration it builds, and the creativity it bring

machine learningaigo
View job →
O
OpenAI
📍 San Francisco• Full-time
17 days ago

About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,

awskuberneteslinux
View job →
O
25 days ago

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We’re looking for a Robotics Control Systems Engineer to take on a foundational role within our robotics team. You’ll help architect, implement, tune, and verify the control infrastructure that enables intelligent, reliable, and responsive robot behavior. This is a deeply hands-on role focused on real-time systems, actuation, dynamics, low-level hardware interaction, and whole-robot performance. You’ll spend significant time working directly with robots onsite: debugging behavior, tuning subsystems, running experiments, and providing feedback across mechanical, electrical, and software teams. This role is based in San Francisco, CA, and requires in-person 5 days a week. In this role, you will: Design and implement real-time control algorithms for robotic systems, including motion control, feedback loops, state estimation, actuator control, and subsystem tuning. Define the control architecture from low-level actuators and hardware interfaces through whole-robot behavior and policy. Identify and characterize actuator, hardware, and software parameters through rigorous experimentation, testing, commissioning, and verification. Work with machine learning engineers to implement reinforcement learning models. Collaborate across mechanical, electrical, and software teams to integrate control logic with sensing and actuation hardware. Help inform the mechanical and electrical design to maximize capability and flexibility. Create the control system architecture; determine the correct level of abstraction from actuators all the way up to whole-robot

awsrestmachine learning
View job →
🔔

Get new model behavior engineer jobs by email

Daily job updates · Unsubscribe anytime