About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Revenue team plays a critical role in enabling OpenAI to scale commercial offerings by overseeing billing operations, deal desk, revenue systems, revenue accounting, and controllership. We work cross-functionally with Product, Engineering, Go-To-Market, Tax, Legal, and Technical Accounting to support new monetization strategies, improve operational efficiency, and maintain financial integrity as the business grows. About the Role This senior leader will own key elements of Ads revenue accounting from technical assessment through operational execution. The role will guide accounting for products, pricing, contracts, incentives, credits, refunds, makegoods, international expansion, and new go-to-market motions. It will establish governance and translate approved accounting positions into launch, billing, data, close, reconciliation, and control requirements. Success requires deep technical revenue expertise, strong business partnership, and the ability to build durable 0-to-1 processes in a fast-changing environment. Advertising is a critical and growing monetization vector for OpenAI, and this role will help shape the financial foundations that enable Ads to scale responsibly and transparently. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead technical accounting assessments for Ads products and commercial arrangements, including performance obligations, variable consideration, allocation, principal-versus-agent, collectibility, contract modifications, refunds, incentives, credits, makegoods, and revenue presentation. Own and continuously evolve Ads revenue accounting policies and operating guidance as product behavior, pricing, contracting, incentive programs, and billing models change. Establish governance for new pro
Jobiba hiring network
Model Behavior Engineer Jobs
4,989 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current model behavior engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team The Future of Computing Research team is an applied research team in the Consumer Devices group focused on developing new methods and models to support our vision as we advance forward in our mission of building AGI that benefits all of humanity. About the Role As a Technical Lead on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. This role is based in San Francisco, CA. We follow a hybrid model with 3 days a week in the office and offer relocation assistance to new employees. In this role, you will: Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models. Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth. Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities. Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads. Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems. Run through the necessary walls to take nascent research capabilities and turn them into capabilities we can build on top of. You might thrive in this role if you: Have experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators. Understand the performance characteristics of transformer models, including attention, KV-cache behavior, and memory bandwidth requirements. Have designed or optimized high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines. Have experience building or leading teams work
About the Team The Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously. Our work spans three areas: Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks. Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work. Oversight : Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful agent autonomy (for example future versions of auto-review ). About the Role This role focuses on oversight and system-level mitigations that enable increasingly capable agents to operate safely and autonomously in real environments. We prioritize building oversight systems that are used in practice today, both internally and externally (see our recent work on action monitoring for codex and former code review ). We also study longer-term questions about how increasingly capable agentis systems can be supervised, constrained, and corrected. We’re looking for a safety&security minded researcher or engineer who can reason rigorously about security boundaries and agent behavior, then build and test practical mitigations. A background in AI control or security is welcome but not required. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and evaluate system-level controls for agent actions like agent-based review. Plan how they fit in a broader syste
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . Payments Risk Analyst II As a Payments Risk Analyst II on the Risk & Fraud team within the Consumer & Business group, you'll own the identification and mitigation of fraud risk across Coinbase's fiat payment rails, protecting the company's bottom line and millions of customers without degrading the user experience. You'll join a team that is modernizing its risk controls and building next-generation fraud prevention systems, giving you the opportunity to shape how we protect customers at scale across one of crypto's highest-volume retail platforms. What you'll do: Own end-to-end risk controls for ACH payments, scaling fraud prevention processes to keep pace with business expansion. Drive detection and mitigation of fraudulent deposit activity by implementing targeting logic against bad actors and high-risk behavior. Lead quantitative analysis of fraud performance, measuring rule precision and recall, assessing the financial impact of decisions, and surfacing tradeoffs to leadership with clear, data-backed recommendations. Partner cross-functionally with Product, Growth, Data Science, Risk ML, and Engineering to develop and iterate on risk management strategies, including the transition from heuristic rules to model-driven and synchronous friction controls. Execute investigations into anomalous customer and transaction behavior, proactively identifying
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might Design and run experiments th
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? This role will focus on evaluating coding tasks, requiring you to review and debug code, navigate repository architecture, and analyze model trajectories. Your work will contribute to our model development efforts and the logic our models apply when completing task requests. Please note : This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included! As a Data Annotation Specialist, you will: Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools. Assess agent trajectories and model capabilities for code generation and debugging requests. Prompt models to complete complex coding tasks and review the accuracy of generated responses. Label, proofread, and improve machine-written and human-written software engineering-related outputs. Report quality and performance trends related to model/agent behavio
Lead Systems Engineer - Phantom Works Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Lead Systems Engineer (Level 4) to support the Space Battle Management, Command and Control (SBMC2) programs in Colorado Springs, CO or Berkeley, MO . The Systems Engineer will perform as part of a high-performing team and have the opportunity to contribute to the development and delivery of innovative software solutions in an agile software development environment interfacing with internal and external stakeholders; to include a Joint Industry partnership Team (JIPT) and Working Groups, contributing to the design and development of next generation ground command and control capabilities. As a member of the Boeing team, you will be responsible for key portions of our development lifecycle, from idea creation and development, all the way through to maintenance and support of the customer’s delivered system. More importantly, you will have the opportunity to make an impact on the results of our projects. We offer a collaborative mentoring environment where you have the opportunity to learn from others and be a mentor to others. Position Responsibilities Support definition of requirements, interfaces, and concept of operations through supporting and/or chairing working group meetings. Understand and communicate prioritization of efforts to multi-disciplinary team. Think abstractly and see the big picture while evaluating technical details Perform technical analyses to develop and validate models of system behavior; identify solutions to complex problems. Serve as a direct interface to both internal and external customers Conduct and support trade s
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for a Detection & Response Engineer to build the systems that help us identify, investigate, and respond to threats across our platform. This is an engineering role focused on automation. You'll build detections, investigation tooling, and response capabilities that scale with our infrastructure, using AI where it meaningfully improves signal, investigation speed, and operational effectiveness. You'll work closely with infrastructure, platform, and security engineers to ensure every incident makes the platform more resilient. What You'll Work On: Detection Engineering Design and build high-fidelity detections for attacks, abuse, and anomalous behavior across our infrastructure and production systems Continuously improve detections based on telemetry, threat intelligence, and lessons learned from incidents Improve visibility across cloud infrastruc
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE The largest, most demanding enterprises run on Baseten, and they bring exacting requirements for how people, services, and agents access the platform. This is the founding role for our identity and authorization team within enterprise engineering. You'll own the identity and access layer of the Baseten platform: the authorization model, credential systems, and admin experiences that enterprise IT teams use to govern access for organizations like Harvey, HubSpot, and Notion. You'll design and build Baseten's fine-grained authorization system from the ground up to support the workflows customers depend on today while giving them cleaner, more precise ways to manage access as the platform grows. Authorization at Baseten requires low-latency permission checks at high request volume, consistent contracts and behaviors across the product suite, and strong security guarantees for mission-critical, highly regulated workloads. EXAMPLE INITIATIVES Recent and upcoming work in this area: Fine-grained authorization for users, service accounts, and agentic workloads: per-resource permissions at the organization, team, and workload scope to support both common workflows and complex enterprise access policies Programmatic authentication allowing high-compliance customers to connect service principles securely via short-lived, workload-based credentials Agent credentials that grant an agent exactly the access it needs for the gi
We're looking for a Senior Platform Reliability Engineer who brings strong software engineering skills and a deep understanding of system behavior under load and stress. This role is a good fit for someone who wants to own reliability as a first-class concern – building the foundational systems that protect Asana's platform, not just responding when things go wrong. You'll build core platform systems like load shedding, rate limiting, circuit breakers, and traffic controls that protect Asana under real-world load. This is deep, cross-cutting work that shapes stability and performance of our entire infrastructure – and you'll partner closely with other platform teams to make reliability something that's built in, not bolted on. Our tech stack includes: AWS, Kubernetes (EKS), CloudFront, Istio, Cilium, MySQL (RDS), OpenSearch, DynamoDB, Redis, Terraform, Datadog, TypeScript, Scala, Go, and Python. (Yeah, we know this sounds like buzzword bingo – but we want this post to actually show up in your searches.) Why this role? Reliability as a first-class feature : You won't be patching things up after the fact. You'll build the systems that make Asana resilient by design. Foundational work : Load shedding, traffic management, ingress/egress – these are the building blocks that protect everything else. You'll own them. Strong collaboration, reasonable hours : You'll work closely with infrastructure teams in Warsaw and Reykjavik, making deep collaboration practical without constant timezone gymnastics. Room to grow : This is a new team, and you'll help shape what Platform Reliability Engineering looks like at Asana – whether that means leading projects, mentoring others, or defining our technical direction. In this role, success means shipping systems that other teams rely on by default – because they make the platform safer, not because they're mandatory. We're especially interested in people who think like backend engineers but obsess over failure modes, capacity plan
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Luxury team builds rich and creative features that set the standard for the ride sharing industry. We are looking for a motivated software engineer to join us building features for Black, Black SUV, and other upmarket modes where quality is uncompromising, where we work on balancing and growing these unique marketplaces within the wider Lyft ecosystem and influencing driver behavior to facilitate highest-quality rides. Ownership is a key quality for this team; this person should be driven to track a project to successful completion and beyond, taking initiative to work with other teams and functions to ensure the code they write reaches users and drives impact. You'll collaborate with engineering, product, data science, analytics, and operations on programs that empower us to iterate quickly, delighting our passengers and drivers. Preferred applicants intend to work regularly from our CDMX office, where team members collaborate together in person. Responsibilities: Help establish roadmap and architecture based on technology and understanding of customer needs Write well-crafted, well-tested, readable, maintainable code Participate in code reviews to ensure code quality and distribute knowledge Share your knowledge by giving brown bags, tech talks, and promoting appropriate tech and engineering best practices Can help lead large projects from idea to positive execution Unblock, support and communicate with internal partners to achieve results Experience: Engineering industry experience Experience with object-oriented programming Experience in distributed systems Experience working with databases, relational or NoSQL Write clear, scalable and clear design documentation Design, build and improve a set of team owned components Please submit your resume in English.
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
AI Engineer - Enterprise Search Overview We are looking for an experienced Enterprise Search Lead to build and optimize our multi-tenant enterprise search solution. This role focuses on creating scalable systems that integrate seamlessly with customer environments, leveraging cutting-edge AI and ML technologies. You will design and manage the enterprise knowledge graph, implement personalized search experiences, and drive AI- powered innovations to enhance search relevance and ITSM workflows. What You Will Do: Build and structure our enterprise knowledge graph to organize content, people, and activity into meaningful relationships for better search relevance. Develop and refine personalized ranking models that adapt to user behavior and improve search results over time. Design ways to adapt AI language models to each customerʼs data for enhanced accuracy and context. Explore innovative methods to combine LLMs with search engines for answering complex queries Write clean, robust, and maintainable code that integrates smoothly with multi-tenant systems. Collaborate with cross-functional teams to align search capabilities with ITSM workflows. Mentor junior engineers or learn from experienced ones to grow as a technical leader. What You Should Have 3-5 years of experience working on enterprise search products with AI/ML integration. Expertise in multi-tenant systems and securely integrating with external customer systems. Hands-on experience with tools like Elasticsearch, Solr, or similar search platforms. Strong coding skills in Python, Java, or equivalent languages. A passion for solving complex problems with AI and delivering intuitive user experiences. Important notice for candidates: Job scams are on the rise. Please keep these guidelines in mind when applying for any open roles at Atomicwork. Only apply through official Atomicwork channels. We do not use third-party agencies or individuals who ask for payments in exchange for interviews or offer letter
Staff Engineer - Enterprise Search Overview We are looking for an experienced Enterprise Search Lead to build and optimize our multi-tenant enterprise search solution. This role focuses on creating scalable systems that integrate seamlessly with customer environments, leveraging cutting-edge AI and ML technologies. You will design and manage the enterprise knowledge graph, implement personalized search experiences, and drive AI-powered innovations to enhance search relevance and ITSM workflows. What You Will Do Build and structure our enterprise knowledge graph to organize content, people, and activity into meaningful relationships for better search relevance. Develop and refine personalized ranking models that adapt to user behavior and improve search results over time. Design ways to adapt AI language models to each customer’s data for enhanced accuracy and context. Explore innovative methods to combine LLMs with search engines for answering complex queries. Write clean, robust, and maintainable code that integrates smoothly with multi-tenant systems. Collaborate with cross-functional teams to align search capabilities with ITSM workflows. Mentor junior engineers or learn from experienced ones to grow as a technical leader. What You Should Have 5+ years of experience working on enterprise search products with AI/ML integration. Expertise in multi-tenant systems and securely integrating with external customer systems. Hands-on experience with tools like Elasticsearch, Solr, or similar search platforms. Strong coding skills in Python, Java, or equivalent languages. A passion for solving complex problems with AI and delivering intuitive user experiences. Important notice for candidates: Job scams are on the rise. Please keep these guidelines in mind when applying for any open roles at Atomicwork. Only apply through official Atomicwork channels. We do not use third-party agencies or individuals wh
Get new model behavior engineer jobs by email
Daily job updates · Unsubscribe anytime