Jobiba hiring network

Model Behavior Engineer Jobs

4,989 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current model behavior engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

E(
Ema (Enterprise Machine Assistant)
📍 San Francisco Bay Area• Full-time
1mo ago

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The residency You own one hard problem end to end. You write the proposal, build the system, design the evaluation, ship behind a gate, and finish with a write-up of what turned out to be true, including the parts that didn't work. You'll sit in the production codebase with a senior mentor and real production data. Recent residents have shipped self-improving harnesses, inference-cost work, agent memory, and eval infrastructure. Your project gets scoped with you, not handed to you. The problem space The loop we care about: production traces become data, data becomes training and evaluation, and better agents produce better traces. Projects live somewhere on that loop. Harness and inference-time work. Context engineering, tool and skill design, orchestration, and deciding where extra inference compute actually pays. Self-improvement loops run behind hard fences. Post-training for agents. SFT on curated trajectories, preference optimization, RL on real agent tasks. Reward design where outcomes are verifiable, process vs. outcome supervision, distilling frontier behavior into cheaper models. Environments and rewards. Turning enterprise workflows into training and eval environments: fi

pythonaigo
View job →
TI
TEGNA India
📍 Chennai• Full-time
16 days ago

TEGNA Inc. helps people thrive in their local communities by providing the trusted local news and services that matter most. With 64 television stations in 51 U.S. markets, TEGNA reaches more than 100 million people monthly across web, mobile apps, streaming, and linear television, while also maintaining a strong global presence in India with offices in Bangalore and Chennai that support technology, product, and business operations initiatives. Together, we are building a sustainable future for local news. Position Overview TEGNA is looking for a Senior Data Analyst is seeking a skilled and forward-thinking CloudEngineer to join our growing technology team. We are a demand-side platform (DSP) helping advertisers and agencies programmatically reach their target audiences at scale. We are looking for a Senior Data Analyst to drive insights across advertiser performance, audience quality, and supply-side partnerships — going beyond surface-level reporting to uncover the stories that data tells. In this role, you will work at the intersection of data, engineering, and data science — collaborating closely with our engineering team to maintain data integrity and with our ML team to translate model outputs into business intelligence. You will analyse how audiences behave across our platform, evaluate the quality and efficiency of inventory sources, develop forecasts that help advertisers and internal teams plan with confidence, and surface actionable intelligence that helps our advertisers deliver better outcomes while growing the overall business. What You’ll Do • Analyse audience behaviour and segment performance to help advertisers improve targeting efficiency, reach, and campaign ROI across programmatic channels. • Evaluate supply-side inventory quality — assessing publisher segments, bid stream data, win rates, and CPM trends — to inform smarter buying decisions and supply curation strategi

sqlgitai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

pythonawsgit
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. At Micron Technology, we transform how the world uses information to enrich life for all. The Heterogeneous Integration Group (HIG) HBM Architecture team develops next-generation High-Bandwidth Memory (HBM) solutions that power AI, high-performance computing, cloud infrastructure, and advanced networking systems. The team works across architecture, design, verification, packaging, product engineering, and technology development to evaluate innovative memory architectures and deliver scalable, high-performance semiconductor solutions. As an HBM Design Architect, New College Graduate, you will contribute to the evaluation and development of future HBM and DRAM architectures. Working with experienced architects and engineering teams, you will analyze system and block-level design tradeoffs related to performance, power, area, thermal behavior, reliability, and manufacturability. This role provides an opportunity to leverage AI, Large Language Models (LLMs), and data-driven engineering methodologies to accelerate architecture exploration and improve decision-making. Responsibilities Analyze HBM and DRAM architectures using analytic

airecruitment
View job →
C
Cohere
📍 European Union• Full-time• Remote
14 days ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About The Role Help bring Cohere's enablement strategy to life for our field teams in EMEA. This generalist role will build and manage programs end-to-end, from design through rollout, measurement, and iteration, while working directly with managers and reps to ensure enablement drives results. You'll report to the Head of Sales Enablement and help establish operational rigor as this global function scales. Key Responsibilities Own enablement programs from design through rollout, measurement, and iteration Partner with Product, Marketing, Sales Leadership, and RevOps to align GTM priorities and embed enablement into every stage of the sales lifecycle Create and curate high-impact resources (playbooks, battle cards, demo scripts, certification programs) for Cohere's solutions Establish KPIs (win rates, sales cycle efficiency) and use analytics to refine initiatives and prove ROI Operationalize MEDDPICC methodology into everyday seller behavior (deal reviews, coaching, planning) Build certification paths that accelerate new seller productivity across regions and cultures Maintain and evolve the enablement content platform to ensur

REMOTEgitaigo
View job →
C
Cohere
📍 San Francisco• Full-time• Remote
18 days ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About The Role Help bring Cohere's enablement strategy to life for our field teams in North America. This generalist role will build and manage programs end-to-end, from design through rollout, measurement, and iteration, while working directly with managers and reps to ensure enablement drives results. You'll report to the Head of Sales Enablement and help establish operational rigor as this global function scales. Key Responsibilities Own enablement programs from design through rollout, measurement, and iteration Partner with Product, Marketing, Sales Leadership, and RevOps to align GTM priorities and embed enablement into every stage of the sales lifecycle Create and curate high-impact resources (playbooks, battle cards, demo scripts, certification programs) for Cohere's solutions Establish KPIs (win rates, sales cycle efficiency) and use analytics to refine initiatives and prove ROI Operationalize MEDDPICC methodology into everyday seller behavior (deal reviews, coaching, planning) Build certification paths that accelerate new seller productivity across regions and cultures Maintain and evolve the enablement content platform

REMOTEgitaigo
View job →
A
Airbnb
📍 United States• Full-time• From $200K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Airbnb is a vision and mission driven company, and our Product Managers embody that mindset. Platform PMs imagine the ideal end state for our community first, work backwards from it, and deliver it in a scalable way. The AI Assistant team owns the agentic AI system that powers Airbnb's support experience for millions of guests and hosts. This is one of the highest-visibility applied-AI efforts at the company, and it sits at the intersection of large language models, platform thinking, and a genuinely two-sided marketplace. The Difference You Will Make: You will own the platform that determines how our AI Assistant reasons, retrieves, and responds: the layer that interprets each request and routes it, the knowledge and capabilities the assistant draws on, the actions it can take to actually resolve an issue, and the evaluation systems that keep it safe and accurate at scale. You will own what a good outcome looks like for the user, and the criteria we measure it against, working through partners who own the underlying knowledge and the engineering implementation. This is a role where Product both directs technical work and does it themselves. You will diagnose architectural problems, author the artifacts that become production behavior, and help drive engineering and data science decisions alongside your key partners. Support is high-stakes: people reach out when something has gone wrong, often with another party involved. You will be responsible for making those moments accurate, safe, and genuinely helpful, increasing how often the assistant fully and correctly resolves

S
Stripe
📍 Seattle• Full-time• $192K – $288K/yr
1mo ago

Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Build statistical models, define and analyze product and operational metrics, explore experimental design, and construct exploratory analysis with internal data. Work closely with product and business teams to identify important questions and answer them with data. Collaborate with other data scientists, engineers and operations to formulate innovative solutions to experiment and implement advanced data mining techniques. Conduct exploratory analysis on internal data to understand user behavior to inform product development. Drive the collection of new data and the refinement of existing data sources. Apply statistical and machine learning models on large datasets to measure results and outcomes, and identify causal impact and attribution. Predict future performance of users or products. Define, measure, and monitor key outcome metrics for teams and support Stripe’s business. Communicate complex concepts and the results of metrics and analyses in a clear and effective manner through creative visualization. Communicate findings broadly and interact with other teams including product managers, software engineers, marketing, and business development. Who you are Minimum requirements Must have a Master's degree or foreign equivalent in Operations Research, Statistics, Industrial Engineering, Business Analytics, Mathematics or a related field, plus three (3) years of

pythonsqlmachine learning
View job →

About the Team The GTM Data Science team partners with Go-to-Market, Technical Success, Product, Engineering, RevOps, and Strategic Finance to build the shared intelligence layer for OpenAI's B2B business. The team turns product usage, customer behavior, revenue, field activity, and customer feedback into rigorous insight products that help leaders and field teams understand where customers are succeeding, where adoption is blocked, and what actions will accelerate durable growth. We are building systems that make customer intelligence proactive: surfacing risk, expansion potential, product gaps, and repeatable playbooks before they show up as escalations or missed opportunities. About the Role As the Applied Data Science & Insights Lead for GTM Intelligence Solutions and Technical Success, you will be a hands-on technical leader responsible for shaping how OpenAI measures, understands, and improves customer adoption across our B2B products. You will build AI/ML-powered intelligence products that connect account health, product usage, customer lifecycle, support tier, qualitative sentiment, commercial context, and field actions into a practical operating system for GTM and Technical Success. This role will build the data science foundation for Technical Success: defining the metrics, models, operating insights, and decision systems that help the team scale customer adoption and expansion with rigor. You will also be expected to build and lead a small mighty team over time: setting direction, hiring and developing talent, creating operating cadences, and holding a high bar for technical rigor and business impact. You will lead the development of models, metrics, and decision systems that recommend what GTM and Technical Success teams should do next, explain why, and measure whether those interventions worked. Your work will help customers move from pilots to production, deepen usage across products, identify high-value use cases, reduce churn risk, and create a f

pythonsqlaws
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $216K/yr
16 days ago

Scale Labs, Research Scientist — AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on AI Controls and Monitoring, you will design methods, systems, and experiments to ensure that advanced AI models and agents remain aligned with intended goals, even in high-stakes or adversarial environments. For example, you might: Develop monitoring techniques and observability methods that track AI behavior in real time to identify and flag deviations, emergent capabilities, or anomalous outputs; Research mechanisms for layered control, including fail-safes, oversight protocols, and intervention methods that can halt or redirect AI systems when risks are detected; Design red-team simulations to probe weaknesses in oversight and control mechanisms, and build mitigations to close identified gaps; Collaborate with policymakers, engineers, and other researchers to establish standards and benchmarks for AI monitoring and escalation. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be

awsrestmachine learning
View job →

Here's a summary of the role Data is only powerful when everyone agrees what it means. This is your chance to build a world class product behavioural data system and process from the ground up. We're rebuilding the internal product our teams use to measure product success and end-user value across all our 24+ products in four business units. You'll design the single source of truth: the metric definitions, the telemetry standards, the data contract engineering instruments against, the change governance processes and tooling, and the models that turn behaviour into portfolio insight. You'll educate and support the teams using this internal behavioural data product to set expectations and and plot a rational roadmap for the maturing of this product, helping our user to use the data honestly and wisely. Usage tells you what happened, not why, or whether it mattered. We want someone who sets quantitative behaviour against qualitative evidence and treats the disagreement as the interesting part. AWS/Snowflake is our spine; capture and BI are open decisions you'll help make. Your models feed the business reviews our executive team, CTO and investors use to run the portfolio, and engineering is committed to instrument against your contract. Here's what your first twelve months should look like First 90 days. Learn the estate, meet the engineers who instrument it, and agree the measurement model for one product end to end — with a defensible adoption number and the qualitative read on what it means. Three to six months. Data contract published, instrumentation spec landed with engineering, first business unit building against it. Definitions, naming and versioned change control governed. Six to twelve months. Rolling out across the remaining business units through the BU-aligned analysts, portfolio reporting running on your numbers. Here's a breakdown of what you'll do (not all

sqlawsgit
View job →
O
1mo ago

About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review ). About the Role We’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity. Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals. Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems. You might thrive in this role if you: Have demonstrated strength in research engineering, ML en

awsrestai
View job →
B
Baseten
📍 San Francisco• Full-time• Remote
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption

REMOTEpythonsqlmachine learning
View job →
N
Notion
📍 San Francisco• Full-time• $196K – $230K/yr
1mo ago

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for the end-to-end product experience where people discover, set goals, delegate work, review results, and build trust over time with AI. This role sits at the intersection of research craft and evaluation operations: you’ll run studies that uncover user mental models, expectations, and failure/recovery behaviors, then translate those insights into reusable rubrics, workflows, and measurement approaches that product, design, engineering, and data science can apply consistently. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Define what “good” looks like (frameworks & rubrics): Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness,

pythonsqlgit
View job →
Z
Zscaler
📍 Bellevue• Full-time• From $180K/yr
16 days ago

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Senior Staff Rust Developer to join our Platform Convergence Team. This is a hybrid role based in San Jose, CA reporting to the Sr. Director, Software Engineering. Join us to build a new platform from the ground up that can scale hundreds of millions of users with high reliability and low latency. You will design and implement distributed system and core infrastructure components while collaborating closely with various stakeholders. What you’ll do (Role Expectations) Design and build a low-latency, high-throughput data forwarding plane using Rust, leveraging its async/await model for efficient I/O and service-oriented infrastructure Develop distributed, scalable systems with a focus on concurrency, fault tolerance, and messaging Implement and maintain gRPC-based APIs and services to integrate forwarding plane capabilities with control and orchestration layers Optimize system

awskubernetesci/cd
View job →
🔔

Get new model behavior engineer jobs by email

Daily job updates · Unsubscribe anytime