The Agent Research and Tooling team, part of MongoDB's AI Builder Experience organization, owns the platform layer around agents: how teams author, distribute, evaluate, monitor, and improve agent skills and agent behavior. We are hiring a software engineer to build and maintain the tooling, evaluation systems, and quality gates behind MongoDB's agent skills. This is a software engineering role at the intersection of developer tooling, applied AI, and software quality. You will take loosely defined agent and tooling problems, break them into workable plans, and ship durable internal systems: command-line tools, reusable libraries, evaluation harnesses, and CI workflows. This role is open to remote work in the US or can be based out of any of our US offices. What you'll do Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI Integrate tooling into GitHub Actions and other CI workflows, including secrets, annotations, exit codes, and artifacts Build code-generation quality checks, such as anti-pattern catalogs and linting for AI-generated MongoDB code Investigate real failures such as nondeterministic results, false positives, and unsafe generated guidance, and turn them into reusable improvements Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams Examples of the problems you'll solve How can tests verify an agent tool's
Jobiba hiring network
Applied Ai Engineer Jobs
764 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current applied ai engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Agent Research and Tooling team, part of MongoDB's AI Builder Experience organization, owns the platform layer around agents: how teams author, distribute, evaluate, monitor, and improve agent skills and agent behavior. We are hiring a software engineer to build and maintain the tooling, evaluation systems, and quality gates behind MongoDB's agent skills. This is a software engineering role at the intersection of developer tooling, applied AI, and software quality. You will take loosely defined agent and tooling problems, break them into workable plans, and ship durable internal systems: command-line tools, reusable libraries, evaluation harnesses, and CI workflows. This role is open to remote work in Canada or can be based out of any of our Canada offices. What you'll do Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI Integrate tooling into GitHub Actions and other CI workflows, including secrets, annotations, exit codes, and artifacts Build code-generation quality checks, such as anti-pattern catalogs and linting for AI-generated MongoDB code Investigate real failures such as nondeterministic results, false positives, and unsafe generated guidance, and turn them into reusable improvements Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams Examples of the problems you'll solve How can tests verify an agent to
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the role: Roblox Studio is the creation engine behind millions of immersive 3D games built by creators around the world. We are entering a major platform transition, evolving Studio into an AI-Native IDE where intelligent systems plan, act, validate, and iterate alongside creators to amplify their output. We are seeking a Technical Director, Applied AI to own the technical direction of the agentic AI systems embedded deeply into Roblox Studio. This is a hands-on individual-contributor role operating at the intersection of platform engineering, AI systems, and developer experience, where you will set the technical direction, design the architecture, and write the code. You will: Design and build the agentic AI systems at the core of the AI-Native IDE (planning, execution, validation, evaluation) so they are production-grade and reliable. Decide how Studio uses coding models, agents, retrieval, tool calling, and context management across the Assistant and the broader creator workflow. Own the quality bar and evaluation strategy , standing up quantitative and qualitative eval pipelines, including human-in-the-loop, so we ship with confidence. Optimize these systems for latency, throughpu
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With a billion rides per year and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Rider, Marketplace, Growth, and beyond. While traditional approaches to optimization and problem decomposition are sufficient to disrupt transportation, building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives warrants modern ML utilizing peta-byte scale data. Our highly motivated Machine Learning Engineers work on these challenging problems and define solutions to directly impact various aspects of our core business. If you are a critical thinker with experience in machine learning workflows and LLMs, passionate about solving business problems using data and working in a dynamic, creative, and collaborative environment, we are searching for you. We are seeking a Senior Machine Learning Engineer to join the Rider Applied AI team and lead the design, development, and deployment of state-of-the-art machine learning and artificial intelligence systems. This role requires a strategic thinker who can balance high-level system architecture with hands-on technical implementation. You will collaborate across teams to shape the future of ride-sharing by leveraging AI, Machine learning and Data science. Responsibilities: Model Development & Research: Design, build, and deploy machine learning models for real-time applications, including translating state-of-the-art research into production-ready solutions. System Design: Architect scalable, reliable ML pipelines that integrate seamlessly with existing backend systems. Innovation & Applied Research: Stay ahead of the curve by exploring emerging algorithms, technologies (such as LLMs and LLM-based applications), and frameworks — critically eva
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With a billion rides per year and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Rider, Marketplace, Growth, and beyond. While traditional approaches to optimization and problem decomposition are sufficient to disrupt transportation, building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives warrants modern ML utilizing peta-byte scale data. Our highly motivated Machine Learning Engineers work on these challenging problems and define solutions to directly impact various aspects of our core business. If you are a critical thinker with experience in machine learning workflows and LLMs, passionate about solving business problems using data and working in a dynamic, creative, and collaborative environment, we are searching for you. We are seeking a Senior Machine Learning Engineer to join the Rider Applied AI team and lead the design, development, and deployment of state-of-the-art machine learning and artificial intelligence systems. This role requires a strategic thinker who can balance high-level system architecture with hands-on technical implementation. You will collaborate across teams to shape the future of ride-sharing by leveraging AI, Machine learning and Data science. Responsibilities: Model Development & Research: Design, build, and deploy machine learning models for real-time applications, including translating state-of-the-art research into production-ready solutions. System Design: Architect scalable, reliable ML pipelines that integrate seamlessly with existing backend systems. Innovation & Applied Research: Stay ahead of the curve by exploring emerging algorithms, technologies (such as LLMs and LLM-based applications), and frameworks — critically eva
Sendbird is building AI agents for customer experience. Our platform already powers billions of conversations every month across chat, voice, video, and messaging APIs. We are now using that foundation to build agents that understand customer context, reason over business data, and take reliable action in production. We are looking for a Machine Learning Engineer to research, build, and productionize new capabilities for those agents. This role sits at the intersection of agent product development, applied AI research, and production engineering. You will work on systems that enterprise customers depend on every day, not demos or isolated prototypes. About Sendbird and delight.ai Sendbird has spent more than a decade building communication infrastructure for in-app chat, voice, video, and messaging APIs. More than 4,000 brands use our platform, including DoorDash, Match Group, Noom, Yahoo Sports, and Rakuten. Our systems support more than 7 billion messages every month. In 2024, we made a strategic shift toward AI-first customer experience. In 2025, we launched our enterprise AI agent product, delight.ai. Delight.ai helps businesses deliver customer support and engagement that is faster, more contextual, and more personal. Unlike simple FAQ bots, our agents are built to remember customer context, use tools, retrieve relevant knowledge, connect across channels, and handle real customer workflows with accuracy and control. The Role As a Machine Learning Engineer, you will design, build, evaluate, and ship new capabilities for our AI agents. You will work across agent architecture, retrieval, memory, planning, tool use, workflow automation, voice, evaluation, data pipelines, model adaptation, inference, and production integration. This is a hands-on engineering role for someone who can turn AI research and product ideas into reliable customer-facing features. Some problems will require training, fine-tuning, or adapting models. Others will require better retrieval, bet
Role Description As a Principal Engineer at Dropbox, you will own company critical, loosely defined technical problems with multi year impact, operating at the intersection of technology, product, business strategy, and applied AI. You will define long term technical direction for customer facing experiences used by millions, identifying where AI meaningfully improves customer value and translating evolving business context and industry advances into durable, multi area strategies and roadmaps that shape how Dropbox builds, scales, and innovates, while remaining hands on in software development where it provides the greatest leverage. Your influence will span organizations, setting foundational architecture, driving execution standards, and aligning senior technical and product leaders across boundaries. You will lead the responsible introduction and adoption of AI across product capabilities and engineering workflows, bring clarity to the most complex decisions, institutionalize engineering excellence, and contribute directly through critical design, prototyping, and code reviews. In return, you will operate as a trusted technical partner to senior leadership, shape systems and platforms including AI powered foundations that define Dropbox’s future, and act as a company level technical strategist, advancing Dropbox’s mission to create a more enlightened way of working. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here. Responsibilities Own and drive technical outcomes across multiple teams and organizations, delivering company critical customer and business impact at scale. Define long term technical strategy and partner with senior Product and Engineering leaders as the technical owner for the most important company objectives. Tackle the most ambiguous and far reaching technical and product problems, shaping wh
The Applied AI team designs and builds algorithmically driven features in the Datadog app. We work across a range of applications, primarily focusing on analysis on streaming data such as anomaly detection, error outliers and faulty deployment analysis. As an Applied Scientist you will work on building models and algorithms for machine learning powered features within the Datadog platform. You will work closely with our engineering and product partners to explore, build, scale and deliver these features that we incubate within the Applied AI team. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design solutions for our different use cases. You will research and benchmark relevant algorithms to find the best fit for our use-cases Leverage machine learning algorithms and statistical techniques to build new scalable product features Develop, deploy and monitor new and existing features to production Participate in our journal club by reading and presenting the latest academic research papers to the team Explore, analyze and tell the story behind high volumes of data flowing through Datadog systems Maintain and monitor the models, services and infrastructure owned by your team Participate in your team’s on-call rotation Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering, Machine Learning or related scientific field or equivalent experience You have experience working with high-scale systems and datasets including building models, applying machine learning to real business problems, and writing production data pipelines You can explain complex ideas and algorithms to non-technical audiences You care about code simplicity and performa
The Applied AI team designs and builds algorithmically driven features in the Datadog app. We work across a range of applications, primarily focusing on analysis on streaming data such as anomaly detection , error outliers and faulty deployment analysis . As an Applied Scientist you will work on building models and algorithms for machine learning powered features within the Datadog platform. You will work closely with our engineering and product partners to explore, build, scale and deliver these features that we incubate within the Applied AI team. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design solutions for our different use cases. You will research and benchmark relevant algorithms to find the best fit for our use-cases Leverage machine learning algorithms and statistical techniques to build new scalable product features Develop, deploy and monitor new and existing features to production Participate in our journal club by reading and presenting the latest academic research papers to the team Explore, analyze and tell the story behind high volumes of data flowing through Datadog systems Maintain and monitor the models, services and infrastructure owned by your team Participate in your team’s on-call rotation Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering, Machine Learning or related scientific field or equivalent experience You have experience working with high-scale systems and datasets including building models, applying machine learning to real business problems, and writing production data pipelines You can explain complex ideas and algorithms to non-technical audiences You care about code simplicity and performance You are excited to work on
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer for our Frontier Security AI team. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers — products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. In this role, you will lead the design and development of our Agentic Harness and agent evaluation platform, working across product, infrastructure, applied AI, security, and modeling teams to take new capabilities from prototype to dependable customer value. AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. Convert ambiguous reports such as "the agent feels worse" into measurable failure modes, reproducible tests, and durable fixes. Analyze production agent trajectories to identif
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Platform team within the Infrastructure organization. As a senior member of Baseten's Platform Team, you will own the systems that let every engineer at Baseten prove their code works before it reaches production. Our product runs mission-critical AI inference for customers who measure downtime in dollars per second, which means our internal bar for correctness, performance, and failure tolerance has to be exceptional. Your focus is the full testing stack: fast and reliable unit test tooling, integration harnesses that spin up realistic environments on demand, load and performance testing for GPU-backed inference workloads, and resilience testing that deliberately breaks things so our customers never have to find out what happens when a node dies mid-request. This is a builder role with org-wide leverage. You won't be writing tests for other teams — you'll be building the frameworks, harnesses, and feedback loops that make writing good tests the path of least resistance, and you'll set the standards for what "well-tested" means at Baseten. RESPONSIBILITIES Own Baseten's testing strategy end to end — define the standards, the tiers, and the tooling that engineering teams build against. Build and maintain unit, integration, load and performance testing frameworks Design end to end test infrastructure that provisions realistic dependencies
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Job Summary As a Business Systems Engineer on the GTM Engineering team, you will design, build, and operate the automation and AI systems that power ClickUp's Go-To-Market business — spanning core business products and the integration platforms that connect them. This role is AI-native at its core: you won't just maintain existing workflows, you'll actively advance our GTM systems with intelligent agents, LLM-powered automations, and next-generation integration patterns built on MCP, ClickUp Super Agents, and modern iPaaS tooling. You'll sit at the intersection of business systems architecture and applied AI — partnering with Sales, Finance, Revenue Operations, and fellow GTM Systems engineers to eliminate toil, accelerate revenue workflows, and build the automated, AI-augmented infrastructure the company runs on. This is a hands-on engineering role for someone deeply fluent in enterprise business systems, excited about deploying production AI, and committed to genuine ownership of the platforms they build — directly supporting GTMSOE's broader mission of operational excellence across the GTM org. Key Responsibilities AI-Native Automation & Agent Development Design and build AI-powered automations and agentic workflows across the GTM tech stack — including Salesforce, NetSuite, Workato, and MuleSoft — to eliminate manual effort and accelerate business operations. Develop and deploy ClickUp Super Agents and LLM-based automations to automate tasks such as deal data enrichment, quote generation assistance, order validation, revenue recognition triggers, and exception handling in quote-to-cash workflow
Team description At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the platform that connects these agents to Datadog: the MCP Server, the tools and retrieval surfaces agents call into, and — critically — the evaluation systems that tell us whether an agent's experience on Datadog data is actually getting better over time. This role is about that last piece. We're hiring a Staff Applied Scientist to define what "good" means for an Agentic interface at Datadog and to build the measurement systems that make it true. "Good" isn't one number — it spans answer quality, tool-selection accuracy, retrieval relevance, latency, token cost, and end-to-end agent success on real customer workflows. You'll design the evals, build the datasets, define the metrics, and partner with the AI engineers on the team to land the platform that lets every product group at Datadog ship integrations that are demonstrably better release over release. The space is full of open research questions. How do you evaluate an agent end-to-end when the trajectory is non-deterministic? How do you score tool selection when the tool catalog has hundreds of entries and grows weekly? How do you build a measurement system that catches regressions across first-party and third-party agents at once, without each team writing their own harness? If those are the problems you want to spend your time on, come build this with us. Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply. What You’ll Do: Own the evaluation strategy for Datadog's AI agent integrations. Define the metrics — offline and online, quali
Role Description As a Principal Engineer at Dropbox, you will own company critical, loosely defined technical problems with multi year impact, operating at the intersection of technology, product, business strategy, and applied AI. You will define long term technical direction for customer facing experiences used by millions, identifying where AI meaningfully improves customer value and translating evolving business context and industry advances into durable, multi area strategies and roadmaps that shape how Dropbox builds, scales, and innovates, while remaining hands on in software development where it provides the greatest leverage. Your influence will span organizations, setting foundational architecture, driving execution standards, and aligning senior technical and product leaders across boundaries. You will lead the responsible introduction and adoption of AI across product capabilities and engineering workflows, bring clarity to the most complex decisions, institutionalize engineering excellence, and contribute directly through critical design, prototyping, and code reviews. In return, you will operate as a trusted technical partner to senior leadership, shape systems and platforms including AI powered foundations that define Dropbox’s future, and act as a company level technical strategist, advancing Dropbox’s mission to create a more enlightened way of working. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Own and drive technical outcomes across multiple teams and organizations, delivering company critical customer and business impact at scale. Define long term technical strategy and partner with senior Product and Engineering leaders as the technical owner for the most important company objectives. Tackle the most ambiguous and far reaching technical and product problems, shaping w
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join The Quality Platform team is at the heart of Airbnb’s mission to deliver a seamless, high-quality experience for millions of hosts and guests. We don’t just find bugs — we build the systems that prevent them. Our team sits at the intersection of Quality Engineering, Infrastructure, and Applied AI. We are evolving how software quality is built by integrating LLMs, intelligent automation, and data-driven systems into the testing lifecycle. You will join a high-impact group of engineers focused on building AI-powered quality systems that scale across one of the world’s most complex codebases. The Difference You Will Make: As a Mobile Software Engineer, you will be a key contributor to the development of our Quality Platform across both native mobile stacks. You will help build the foundation for how quality is engineered across Airbnb's iOS and Android ecosystems, developing tools and frameworks that enable our mobile platform to scale while keeping developers productive and confident, regardless of which platform they build on. In this role, you will: Build AI-Driven Solutions: Contribute to AI-native agents that automate repetitive testing tasks and provide intelligent feedback to developers on both platforms.Deliver Scalable Infrastructure: Develop and maintain the high-scale platforms and testing environments used daily by the iOS and Android engineering organizations.Promote Engineering Craft: Implement best-in-class mobile patterns and modularity to improve testability and fault-tolerance across both native codebases.Contribute to Operational Excellence: Ensure our automated syste
Get new applied ai engineer jobs by email
Daily job updates · Unsubscribe anytime