Jobs in United States

Model Behavior Engineer in United States

2,174 active opportunities · Updated October 2026

Explore current model behavior engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team The ChatGPT Model Flywheel team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. Team Focus Areas Model Experimentation: Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management. Model Deployment: Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling. Automate capacity management and incorporate platform-wide health monitors. Model Measurement: Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards. Improve end-to-end feedback loops for continual model improvement. Key Partnerships Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API. In this role, you will: Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks. Drive expansion and improvement of multi-tier model experiences. Support and scale self-serve experiment capabilities and automated guardrails. Lead model rollout automation, capacity management, and health monitoring. Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.). You might thrive in this role if you have: Proven experience leading engineering teams in complex, cross-functional environments. Demonstrated success shipping production systems at scale (ideally for AI or large backend services). Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling. Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders. Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus. Why Work With Us Tackle highly impactful technical challenges at the cutting edg

AWSRestAIGo
H
📍 United States· Remote
✓ High-confidence listingCompany trend +310%
Quick readStrong listing-quality and freshness signals

Become a part of our caring community The Associate Vice President, Model & AI Governance, is the enterprise leader responsible for establishing and overseeing the organization’s framework for model governance, artificial intelligence (AI) risk management, responsible AI and AI governance. Reporting to the Chief Audit & Risk Officer, this executive provides independent second-line oversight and challenge of the organization’s use of models, advanced analytics, machine learning, generative AI, and emerging AI technologies. The role establishes the governance, risk-management, control, monitoring and escalation framework necessary to ensure AI models are deployed in a manner that is safe, ethical, transparent, explainable, compliant, secure and aligned with the organization’s mission and risk appetite. The Associate Vice President, Model & AI Governance, partners closely with executive leadership, technology, data and analytics, clinical/business leaders, compliance, legal, privacy, cybersecurity, information security, internal audit, enterprise risk management and other control functions to ensure AI-related risks are identified, assessed, governed, monitored, and appropriately reported. This leader will serve as an advisor to executive management on emerging model and AI risks, while maintaining appropriate independence from the teams developing and deploying models and AI solutions. Key Responsibilities Own and continuously enhance the enterprise model governance framework, including model identification, inventory, classification, risk tiering, development, validation, approval, implementation, monitoring, change management, retirement and documentation. Define model risk appetite, risk taxonomy, minimum control standards, governance requirements and escalation thresholds. Provide effective challenge over model

Machine LearningArtificial IntelligenceAIRecruitment
N
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -87.4%
Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: Notion has been at the cutting edge of AI since before ChatGPT launched. The job of the Model Capabilities team is to keep us there. We own the model layer of Notion AI: integrating frontier models as they ship, keeping inference reliable and economical at scale, and building new capabilities that other teams take advantage of. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: Bring new frontier models into production quickly, making them available for our users and our engineers. Make inference reliable: better error categorization, self-healing retries, and cross-provider failover. Own observability for the model layer, driving down both time to detection and time to fix. Build new model-level capabilities and help product teams adopt them. Act as connective tissue across Notion's AI teams: find the gaps, unblock people, and make sure fixes land with the r

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT infrastructure team is responsible for ensuring that our products can serve rapidly growing demand with the performance, reliability, and quality our users expect. This work sits at the intersection of product demand, model deployment, inference, research, fleet, and capacity. The team translates changing product and model needs into clear capacity decisions and safe, scalable launches. About the Role We are seeking a Technical Program Manager to lead the operating system for Chat capacity and model deployment. You will connect demand forecasting and capacity allocation with model readiness, rollout planning, launch coordination, and post-deployment learning. You will also own mode deployment beyond capacity by working with cross functional teams across research, post-training, inference and product to own mainline model deployment. You will bring structure to constrained-capacity decisions, improve the tooling and mechanisms teams use to prioritize demand, and help new models reach users safely and efficiently. Success requires technical depth, sound judgment under ambiguity, and crisp execution across product, research, infrastructure, and operations teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own cross-functional programs for Chat capacity forecasting, allocation, headroom planning, and constrained-capacity operations. Build durable intake, prioritization, and decision mechanisms that connect product demand and model requirements to available serving capacity. Partner

AWSRestAIRust
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Join NVIDIA's NIM team and be part of an exceptionally ambitious project in Santa Clara, CA! As a Senior Software Engineer, NIM Tools, you will have the remarkable opportunity to build a groundbreaking model customization and deployment lifecycle platform from inception. This isn't just another feature team—you will be defining the structure for a new product surface accessed by ISVs and CSPs internationally. Your work will empower customers to take models from selection through fine-tuning, evaluation, deployment, and compliance flawlessly. What you'll be doing: Compose and build the fine-tuning handoff pipeline, including LoRA adapter repackaging, re-quantization, and re-validation into NIM. Develop the evaluation harness, ensuring models meet our high standards. Implement the observability and attestation layer to produce auditable compliance artifacts. Work in close partnership with ISVs and CSPs to roll out NVIDIA NIMs on a large scale. Define and improve durable platform APIs, steering clear of one-off integrations. Ensure flawless completion of projects through strict attention to detail and proven methodologies. Wha

O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability. You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI’s AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production. In this role, you will: Design and implement the LLM inference runtime for frontier models running on custom silicon. Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference. Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization. Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads. Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack. Enable new

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Integrated Marketing team sits at the center of product, brand, creative, media, research, and GTM work. We partner closely with PMM, Product, Developer Relations, Design, Comms, and agency teams to shape launches and campaigns that are clear, distinctive, and grounded in real audience insight. This is a growing function, so the team needs people who can both raise the quality of the work and build the operating model around it: bringing strong judgment, high agency, and trusted partnership to complex, visible moments. About the Role OpenAI business launches move quickly and often involve new products, model capabilities, or research that can change how businesses operate. We’re looking for an integrated marketing manager to help shape those launches from early strategy through execution. You’ll develop a point of view on the audience, positioning, creative direction, and channels, and help turn complex product and research advances into campaigns that connect with customers. You’ll partner across Product, Product Marketing, Research, Brand, Creative, Communications, Growth, and Sales, along with external agencies. The right person understands technology, has strong strategic and creative instincts, and knows how to bring together the right people and ideas to produce work that is clear, credible, and effective. This is a hybrid role based in San Francisco, with three days a week in the office. In this role, you will: Develop integrated marketing strategies for business product launches, new capabilities, model releases, and research moments. Partner with Product, Product Marketing, and Research to understand what’s changing, identify the right audiences, and shape positioning and launch narratives. Lead campaigns that bring launches to life across creative, communications, customer marketing, developer marketing, growth, sales, and paid, owned, and earned channels. Identify customer stories, product demonstrations, and creative concepts that make

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.4%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: We are seeking an experienced Product Marketing Manager with a strong background in engaging developer audiences and delivering impactful go-to-market programs for native AI and enterprise companies. This role requires someone who is both technically savvy and strategic, with a proven track record of crafting compelling product narratives and building marketing assets that resonate with technical decision-makers. This role is specifically focused on our Model API product offering at Baseten. If you’re passionate about AI infrastructure, developer engagement, and simplifying complex technologies for real-world adoption, we want to hear from you. RESPONSIBILITIES: Positioning & Messaging: Develop clear and differentiated messaging that articulates the value of Baseten’s inference platform to developers and enterprise customers. Narrative Development: Shape how the market thinks about closed-to-open weights models and what matters most when building inference. Go-to-Market Strategy: Own the launch process for new features and products, collaborating closely with product, engineering, sales, and growth teams. Content Development: Create high-quality marketing assets, including white papers, technical blogs, demos, and customer case studies. Sales Enablement: Build resources and programs that empower our sales teams to effectively communicate Baseten’s capabilities and benefits. Market Insights: Understand the

Machine LearningAIGo
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may

PythonGitRestMachine Learning
C
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads. As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, e

PythonGitRestAI
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.3%

From $164.7K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a Senior Product Manager for Signal Lifecycle within Trust & Safety, you'll own the product strategy for the ML platform that powers how Pinterest trains, evaluates, deploys, and measures content safety models at scale. You'll lead the development of ML Signal Management — making ML signals first-class entities with unified metadata and identity across systems. Partnering deeply with ML engineering, data science, content safety, and enforcement systems, you'll drive a platform whose scope is expanding from T&S into content quality, ads safety, and beyond. What you'll do: Own and drive the Signal Lifecycle product roadmap, including ML Flywheel infrastructure, auto-deployment, model onboarding, golden dataset management, and signal performance measurement Define and ship ML Signal Management — a unified backbone that elevates ML

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. We focus on performant and efficient model inference, as well as accelerating research progression via model inference. About the Role We are looking for an engineer who wants to take the world's largest and most capable AI models and optimize them for use in a high-volume, low-latency, and high-availability production and research environment. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Work alongside researchers to enable advanced research through awesome engineering. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our model inference stack. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of Azure VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 5 years of professional software engineering experience. Have or can quickly gain familiarity with PyTorch, NVidia GPUs and the software stacks that optimize them (e.g. NCCL, CUDA), as well as HPC technologies such as InfiniBand, MPI, NVLink, etc. Have experience architecting, building, observing, and debugging production distributed systems. Bonus point if worked on performance-critical distributed systems. Have need

AWSAzureRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling, and infrastructure that help improve model performance across OpenAI’s training and inference workloads at frontier scale. About the Role We’re looking for an autonomous, high-ownership developer productivity engineer who cares deeply about helping other engineers move faster, safer, and with more confidence. This role will sit within OpenAI’s Model Performance organization, contributing to developer infrastructure, CI systems, testing workflows, tooling, and broader performance infrastructure efforts. There is also a strong opportunity to contribute to the Triton project and help improve the systems that support performance-critical engineering work across OpenAI. In this role you will: Improve development workflows for engineers working on model performance infrastructure Design and improve CI/CD, release, validation, and testing pipelines Build and maintain tools that improve reliability, iteration speed, and engineering confidence Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows Contribute to infrastructure efforts that support performance-critical training and inference systems Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure Work in a high-context, ambiguous environment where ownership and good judgment matter You might thrive in this role if: You are motivated by enabling the people around you and helping engineers do their best work You have strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows You are highly collaborative, empathetic, and comfortable partnering deeply with technical teams You are strong in Python and enjoy building reliable, scalable developer tools and infrastructure You have experience improving large-scale engineering work

PythonAWSCI/CDRest
P
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity As an Applied Scientist specializing in Small Language Models and AI Training, you will lead research and development efforts focused on building efficient, high-performance language models tailored for practical applications. You will work closely with research, engineering, and product teams to advance model training techniques, optimize architectures, and scale AI solutions. Your work will directly contribute to AI systems that are safe, interpretable, and impactful across diverse usage scenarios. What You’ll Do Lead research and development of novel training methodologies and architectures for small and efficient language models. Design, implement, and evaluate model training experiments to improve performance, robustness, and generalization of language models. Collaborate closely with research scientists and engineers on scalable training pipelines and model deployment strategies. Develop techniques for model compression, fine-tuning, and domain adaptation to optimize models for real-world applications. Ensure AI safety, fairness, and alignment principles are integrated into model training processes and evaluat

PythonMachine LearningAIGo
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building next-generation processors and systems, bringing together world-class expertise across silicon, systems, and software. We are looking for a CPU Performance Modeling Architect to help evaluate, shape, and optimize the performance of future CPU architectures. In this role, you’ll use performance modeling, workload analysis, and deep understanding of CPU architecture to answer complex questions about how a processor should be designed. You’ll work closely with CPU architects, RTL designers, software and compiler teams, and system engineers to identify performance opportunities, evaluate architectural tradeoffs, and turn modeling insights into actionable design decisions. This role is hybrid, based out of Santa Clara, CA or Austin, TX. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A CPU architect, performance architect, or performance modeling engineer with experience influencing CPU architecture or microarchitecture decisions. You have a strong understanding of modern processor architecture and enjoy digging into why a CPU performs the way it does. You are comfortable combining hardware architecture, software, data, and

PythonAWSAIC++
🔔

Get new model behavior engineer jobs in United States by email

Daily job updates · Unsubscribe anytime