At Linear, we're building the product development system for teams and agents. AI is fundamentally changing how software gets built, and we’re shaping the tools this new era requires. Founded in 2019, Linear has become the platform of choice for more than 40,000 companies (including OpenAI, Coinbase, and Ramp) to plan, build, and ship their products. Today, our team is distributed across North America, Europe, and Australia, and we’re continuing to grow internationally. What unites us is relentless focus, fast execution, and a deep care for software craftsmanship. We’re looking for experienced engineers who have shipped applied AI systems to production and want to define what the agent-native future looks like. We are building intelligence into the core of Linear, enabling the product to orchestrate coding, proactively move work forward, and power-up every software team. You’ll work closely with product and design to transform foundation models into structured, reliable workflows embedded deeply in the core of Linear. We care deeply about keeping Linear fast, intuitive, and opinionated—AI is no exception. Location & work mode Linear is a remote-first company, with optional co-working offices in San Francisco, New York, and London. This role is open to candidates based in the North America. You can work from anywhere within this region. We value deep focus and async collaboration, with intentional moments to connect in person through team off-sites, optional co-working, and occasional travel. What you'll do Build AI-powered product features that feel native, fast, and delightful to use Work with product and design to prototype and iterate on intelligent workflows and user interactions Design backend services to power natural language interfaces, smart suggestions, agentic workloads, and more Optimize prompts, fine-tune model behavior, and evaluate performance Help to guide our agent platform, allowing third parties to bring agents into the core Linear experience
Jobiba hiring network
Model Behavior Engineer Jobs
4,989 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current model behavior engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team The Integrity team builds the systems OpenAI uses to understand, prevent, and respond to misuse across our products. We partner with Product, Policy, Safety Systems, User Operations, Security, Legal, Privacy, OpenAI for Government, and research teams to turn policy and threat models into product controls, review workflows, measurement systems, and enforcement paths. About the Role We are hiring a Senior Engineering Leader for Integrity's engineering efforts for Sensitive Deployments where model capabilities, customer requirements, deployment environments, privacy considerations, and misuse risks raise the operating bar significantly. These include hyperscaler and government deployments and other high-stakes environments, regulated and high-trust enterprise settings, zero data retention and privacy-constrained deployments. The workloads are increasingly agentic where harm can emerge across a sequence of actions rather than a single prompt. The Sensitive Deployments Engineering Manager will focus on high-consequence use cases relevant to government deployments and broader deployment-readiness questions for high-risk domains (e.g., healthcare), while building reusable Integrity capabilities for agentic detection and enforcements in these environments. We're looking for a hands-on engineering leader to guide a team of senior engineers working on OpenAI's most critical deployments. They will lead both the people and technical execution of the team: hiring and developing engineers, setting a high technical bar, and driving ambiguous, high-priority initiatives from early problem definition to durable production outcomes. They will partner closely with Product, Research, Security, Global Affairs, Sales, Solutions, and customer technical teams, while going deep on architecture, infrastructure, and model behavior when needed. In this role, you will: Lead and grow a team of engineers responsible for complex, high-impact deployments of OpenAI technology in sensit
About the Team The Human Data team at OpenAI is responsible for identifying and mitigating risks in advanced AI systems by designing evaluations, surfacing vulnerabilities, and collaborating closely with researchers to strengthen model reliability and public trust. About the Role As a Research Program Manager, you will lead initiatives that test the safety and robustness of OpenAI’s models through creative experimentation and structured evaluation. You’ll coordinate efforts across research and engineering teams to transform ambiguous risks into concrete research programs and influence future model development and deployment. We’re looking for people who are technically savvy, comfortable with ambiguity, and excited about shaping the future of safe AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead programs that explore unexpected model behaviors and identify failure modes. Translate vague or emergent risk signals into clear priorities and actionable research plans. Design and run creative evaluations, experiments, and red-teaming campaigns. Collaborate with research, product, and deployment teams to integrate findings into model training and deployment cycles. Develop repeatable systems for tracking model performance and understanding emerging behavior patterns. You might thrive in this role if you: Have strong experience in technical program management, with excellent organizational and communication skills. Are familiar with large language models, prompt engineering, or model evaluation techniques. Are comfortable managing fast-paced, high-uncertainty projects and shaping them from the ground up. Are creative and resourceful in devising new methods for testing model behavior and performance. Can effectively coordinate across technical and non-technical stakeholders to drive alignment and execution. About OpenAI OpenAI is an AI resear
Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities. In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models. You will: Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA. Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Excellent written and verbal communication skills. Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals. Previous experience in a customer facing r
Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling). This role will focus on optimizing data curation and eval to enhance LLM capabilities in both text and multimodal modalities. In this role, you will develop novel methods to improve the alignment and generalization of large-scale generative models. You will collaborate with researchers and engineers to define best practices in data-driven AI development. You will also partner with top foundation model labs to provide both technical and strategic input on the development of the next generation of generative AI models. You will: Research and develop novel post-training techniques, including SFT, RLHF, and reward modeling, to enhance LLM core capabilities in both text and multimodal modalities. Design and experiment new approaches to preference optimization. Analyze model behavior, identify weaknesses, and propose solutions for bias mitigation and model robustness. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning. Excellent written and verbal communication skills Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals Previous experience in a customer facing role. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined du
About the Team The Interpretability team studies internal representations of deep learning models. We are interested in using representations to understand model behavior, and in engineering models to have more understandable representations. We are particularly interested in applying our understanding to ensure the alignment of powerful AI systems. Our working style is collaborative and curiosity-driven. About the Role OpenAI is seeking a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, in close collaboration with a highly motivated team. You will play a critical role in helping OpenAI ensure future models remain safe even as they grow in capability. This will make a significant impact on our goal of building and deploying safe AGI. In this role, you will: Develop and publish research on techniques for understanding representations of deep networks. Engineer infrastructure for studying model internals at scale. Collaborate across teams to work on projects that OpenAI is uniquely suited to pursue. Guide research directions toward demonstrable usefulness and/or long-term scalability. You might thrive in this role if you: Are excited about OpenAI’s mission of ensuring AGI benefits all of humanity, and are aligned with OpenAI’s charter . Show enthusiasm for long-term AI safety & alignment, and have thought deeply about technical paths to safe AGI. Bring experience in the field of AI safety & alignment, mechanistic interpretability, or spiritually related disciplines. Hold a Ph.D. or have research experience in computer science, machine learning, or a related field. Thrive in environments involving large-scale AI systems, and are excited to make use of OpenAI’s unique resources in this area. Possess 2+ years of research engineering experience and proficiency in Python or similar languag
About the Team The Interpretability team studies internal representations of deep learning models. We are interested in using representations to understand model behavior, and in engineering models to have more understandable representations. We are particularly interested in applying our understanding to ensure the safety of powerful AI systems. Our working style is collaborative and curiosity-driven. About the Role OpenAI is seeking a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, in close collaboration with a highly motivated team. You will play a critical role in helping OpenAI ensure future models remain safe even as they grow in capability. This will make a significant impact on our goal of building and deploying safe AGI. In this role, you will: Develop and publish research on techniques for understanding representations of deep networks. Engineer infrastructure for studying model internals at scale. Collaborate across teams to work on projects that OpenAI is uniquely suited to pursue. Guide research directions toward demonstrable usefulness and/or long-term scalability. You might thrive in this role if you: Are excited about OpenAI’s mission of ensuring AGI benefits all of humanity, and are aligned with OpenAI’s charter . Show enthusiasm for long-term AI safety, and have thought deeply about technical paths to safe AGI. Bring experience in the field of AI safety, mechanistic interpretability, or spiritually related disciplines. Hold a Ph.D. or have research experience in computer science, machine learning, or a related field. Thrive in environments involving large-scale AI systems, and are excited to make use of OpenAI’s unique resources in this area. Possess 2+ years of research engineering experience and proficiency in Python or similar languages. Are deeply curious. About OpenA
About the Team OpenAI is building AI systems that can help professionals perform complex, high-value work with greater speed, rigor, and creativity. Investment banking is one of the most demanding environments for knowledge work: bankers must synthesize fragmented information, exercise judgment under pressure, and produce precise, defensible models, analyses, and client materials. Our team works across Research, Product, Engineering, and Go-to-Market to make OpenAI's models genuinely useful for these workflows. We translate real professional work into product requirements, evaluations, training signals, and repeatable customer solutions. We care not only whether a model can generate an answer, but whether it can deliver accurate, defensible work that experienced bankers can trust and use. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. About the Role We are looking for a Subject Matter Expert in Investment Banking to help define what excellent AI-assisted banking work looks like and turn that standard into better models and products. You will bring deep, current knowledge of how investment banking work is actually performed, including company and industry research, financial analysis and modeling, valuation, diligence, transaction execution, and the creation and review of client materials. You will use that expertise to design realistic tasks and evaluations, create and assess high-quality reference work, diagnose model failures, and help our technical teams improve model behavior and product experiences. This is a hands-on individual-contributor role for someone who enjoys both doing the work and explaining what makes it good. You should be comfortable moving between an Excel model, a presentation, a source document, an evaluation rubric, a product prototype, and a conversation with researchers or customers. You will help us distinguish outputs that merely look pl
About the Team The Future of Computing Research team is an applied research team within the Consumer Devices group focused on developing new methods, models, and evaluation frameworks that support our vision for the future of computing. We work at the frontier of multimodal AI, helping turn emerging model capabilities into product experiences that are useful, delightful, and worthy of long-term trust. Our work explores a new class of AI systems that can learn over time, adapt to individuals, and support people in the flow of daily life. This includes long-term memory, user modeling, and personalization systems that are aligned not just with immediate satisfaction, but with a person’s broader goals, values, and well-being. We work closely across research, engineering, design, product, and safety to define what it means to build AI systems that know you over time, act at the right moment, and help in ways that are context-aware, respectful, and demonstrably beneficial. About the Role We are looking for a Research Engineer / Scientist to join the Future of Computing Research team to work on RLHF and post-training for personalized, multimodal AI systems. This role will focus on building the learning and evaluation foundations that help models become more context-aware, adaptive, and useful over time. You will work on problems such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that must make high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded: success is not just higher benchmark performance, but better model behavior in real-world use. The ideal candidate is excited about pushing beyond one-turn assistant behavior toward systems that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value. Internally, that maps closely to the need for careful reward design, feedback loops, and evaluation frameworks that test whether i
About the Team : We build core first-party app experiences in ChatGPT and Codex, define the primitives for high-quality app third-party app experiences, and collaborate with best-in-class partners across consumer and enterprise categories to bring delightful experiences to our customers. About the Role: We are hiring a Product Manager to shape and scale the app ecosystem across ChatGPT and Codex. This person will own both 1P product experiences and partner-led launches, translating user needs, model capabilities, platform constraints, developer and partner requirements, and enterprise controls into products that feel delightful, reliable, and safe. This position is based in San Francisco, CA, with relocation assistance available. In this role, you will: Develop the strategy and roadmap for ChatGPT and Codex app ecosystem experiences across consumer and enterprise use cases. Build and ship high-quality first-party app experiences that demonstrate the best of what apps can do inside ChatGPT and Codex. Collaborate with best-in-class partners to create app experiences that solve real user and business workflows. Define the product foundations and quality standards needed for a trusted app ecosystem. Lead cross-functional execution across engineering, design, research, partnerships, GTM, legal, privacy, security, support, and data/evals. Use customer, user, partner, and model-behavior insights to prioritize the roadmap and improve post-launch performance. You might thrive in this role if you: Have built and scaled consumer or enterprise product experiences with strong product taste and measurable user impact. Have built app platforms, marketplaces, partner ecosystems. Are technically fluent enough to reason about MCPs, APIs, SDKs, and model/product constraints. Can move fluidly between strategy, product, partner judgment, and operational execution. Communicate crisply in writing and bring clarity to ambiguous, fast-moving product areas. Care deeply about user trust, safe
About the Team The Synthetic RL team develops reinforcement learning methods that leverage synthetic data, environments, and feedback to train and evaluate frontier AI models. The team explores approaches such as self-play, simulators, and other synthetic evaluations to push model capability, generalization, and alignment beyond what is possible with the current prevailing methodology. About the Role As a Research Scientist on the Synthetic RL team, you will develop novel reinforcement learning techniques that use synthetic environments and feedback to improve large-scale models. You’ll work closely with other researchers to design experiments, analyze learning dynamics, and translate research insights into training approaches used in production systems. We’re looking for researchers who enjoy working on open-ended problems, value fast iteration, and want their work to directly shape how frontier models are trained. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Research and develop reinforcement learning algorithms Design and run experiments to study training dynamics and model behavior at scale Collaborate with engineers and researchers to integrate successful approaches into model training pipelines You might thrive in this role if you: Have a strong background in reinforcement learning, machine learning research, or related fields Have strong engineering and statistical analysis skills Enjoy exploring new problem spaces where data, objectives, and evaluation are imperfect or evolving Are motivated by seeing research ideas influence real-world AI systems About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an ex
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a CPU Verification Fellow to lead verification strategy and execution for next-generation RISC-V high-performance processors. This role requires deep CPU verification expertise, strong microarchitecture understanding, and the ability to guide large engineering teams from early design through tapeout and post-silicon validation. The ideal candidate has verified complex out-of-order, speculative, superscalar CPUs and can define scalable methodology across simulation, formal verification, emulation, FPGA, and silicon bring-up. This role is hybrid, based out of Santa Clara, CA or Austin, TX. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You have deep experience verifying high-performance superscalar CPUs, ideally including out-of-order and speculative processors. You have strong knowledge of RISC-V architecture, including ISA compliance, privileged architecture, virtual memory, atomics, vector extensions, and memory model behavior. You are highly proficient in SystemVerilog, UVM, constrained-random verification, assertions, functional coverage, and advanced debug methodologies. You have hands-on experience with CPU refere
Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. We are looking for a technically-capable product leader to head up AI Design at Fin. Fin is the #1 AI agent for customer service — a full-stack, vertically integrated AI system powered by our own in-house models, Apex, designed specifically for customer experiences. The role is to lead the Design function within our AI Group, where our custom models and core Fin functionality are developed. You will be managing our AI Designers (currently 3 Staff/Principal level AI Designers). You will work directly with AI Engineers and Scientists in our AI Group , and Product Designers and teams in our R&D group. We believe that leading AI Design is a new type of product leadership role that combines technical, product, and leadership qualities. This includes shaping how the Fin AI models and harness behave — how they reason, respond, and improve over time. The experience is defined as much by model behavior, evals, and iteration loops as it is by UI. We’re looking for someone to lead in that world.
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role We believe that the final enabler for AGI is spending compute on context. As a Context Researcher on Agent Post-Training, you will scale compute spent on context. You will get to work in our frontier training stack on enabling the next paradigm of model training with a clear product interface for iterative deployment (Codex Chronicle). You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments that improve scaling of compute on context. Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. Build evals and environments that expose the next set of model failures,
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of this API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch. This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows. You should be excited to work across research, engineering, data, evals, and product to make models better at acting in real workflows. You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/align
Get new model behavior engineer jobs by email
Daily job updates · Unsubscribe anytime