Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! At Figma we believe design doesn’t end in a file or with your designer - it includes everything that goes into the product you ship: the production code that ties it all together, the systems and developer tools that make that code reliable, the context you provide to our AI agents, and the documentation that keeps everyone on the same page. The Roundtripping Area at Figma is redefining how Designers, PMs and Engineers collaborate. We are responsible for agentic workflows enabling ideation and prototyping on production codebases as well as accelerating the journey from design to code. In 2023, we launched Dev Mode, a suite of features that give developers everything they need to navigate design files and transform designs into code. In 2025, we introduced Figma’s MCP, accelerating how ideas get to production. Looking to the future, we aim to further reduce the barriers between design to code and code to design allowing ideation, prototyping and productionalising to happen seamlessly where it most makes sense. Our Roundtripping team is expanding, and we’re hiring AI Product Engineers across multiple levels in the UK. We’re looking for people who have built generative AI products and are eager to lead AI efforts end-to-end, from early ideas to production. Join us in shaping the future of AI at Figma. What you’ll do at Figma: Build and evolve Dev Mode our MCP tools and Make, Figma’s leading tools for dev/design collaboration Take part in building new 0→1 products within the agentic coding space Col
Jobiba hiring network
Lead Software Engineer Machine Learning Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead software engineer machine learning jobs. Use filters to narrow by work mode, employment type, experience and date posted.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab
About Eudia: Eudia is redefining the future of legal work with AI-powered Augmented Intelligence, enabling Fortune 500 legal teams to move faster, manage risk more effectively, and unlock new business value. Backed by $105M in Series A funding led by General Catalyst, we’re building a category-defining platform that blends AI-driven automation with human expertise, transforming legal from a cost center into a strategic growth driver. At Eudia, we move fast. Unlike traditional enterprise software, our teams ship solutions in days, not months—delivering real impact for some of the world’s largest companies, including Cargill, Coherent, DHL, and Duracell. We’re solving one of the most complex, unsolved challenges in AI: bringing trust, accuracy, and security to legal automation. We’re a team of builders, operators, and problem-solvers who are passionate about reshaping an industry that has long been resistant to change. If you’re looking for a place where you’ll be challenged, take ownership from day one, and work alongside some of the brightest minds in AI and legal —we’d love to meet you. About the Role: Are you interested in building a high-performance Agentic AI driven legal workflow system that supports our current and future scale of platforms? If so, we are looking for you to join our growing team in India. This person will work from our Bangalore office and actively collaborate with the Palo Alto team. We are looking for a Lead Software Engineer that will help develop the most secure, enterprise-grade software using and innovating the latest in Generative AI. The opportunity to tackle challenges in creating cloud-agnostic solutions, maintaining stringent security and compliance standards, and building scalable, resilient platforms for enterprise applications, data, AI, and search also exist while you will be able to routinely innovate on behalf of our customers, collaborating with the world's top
About the Team DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the Role You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly. You’re excited about this opportunity because you will… Lead the design of infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company. Own and evolve our open-weights serving stack — real-time GPU endpoints, high-thr
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe processes over $1.9T in payments volume per year, which is roughly 1.6% of the world's GDP, for millions of customers from startups to enterprises. The tremendous amount of data makes Stripe one of the best places to do machine learning. While being an integral part of almost every product line at Stripe (e.g., Payments, Radar, Capital, Billing, etc.), we have lots of exciting opportunities to innovate in ML Platform at Stripe. The ML Platform team builds the platforms and services that enable ML engineers and data scientists across Stripe to take data and build features and models from prototype to production—reliably, at low latency, and at scale. Our scope spans ML training infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI capabilities. We work closely with product teams, data scientists, and platform infrastructure teams to build powerful, flexible, and user-friendly systems that substantially increase ML velocity across the company. What you'll do You'll serve as a technical lead across the ML Platform space and a key contributor to the evolution of the platforms that power Stripe's ML-driven products. As a Staff Engineer, you'll make decisions with a large impact on Stripe. You'll influence our investments and strategy while making our systems more reliable, secure, and a delight to use. You'll work cross-functionally with other technical staff, dat
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe processes over $1.9T in payments volume per year, which is roughly 1.6% of the world's GDP, for millions of customers from startups to enterprises. The tremendous amount of data makes Stripe one of the best places to do machine learning. While being an integral part of almost every product line at Stripe (e.g., Payments, Radar, Capital, Billing, etc.), we have lots of exciting opportunities to innovate in ML Platform at Stripe. The ML Platform team builds the platforms and services that enable ML engineers and data scientists across Stripe to take data and build features and models from prototype to production—reliably, at low latency, and at scale. Our scope spans ML training infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI capabilities. We work closely with product teams, data scientists, and platform infrastructure teams to build powerful, flexible, and user-friendly systems that substantially increase ML velocity across the company. What you'll do You'll serve as a technical lead across the ML Platform space and a key contributor to the evolution of the platforms that power Stripe's ML-driven products. As a Staff Engineer, you'll make decisions with a large impact on Stripe. You'll influence our investments and strategy while making our systems more reliable, secure, and a delight to use. You'll work cross-functionally with other technical staff, dat
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are We're looking for an innovative and passionate Machine Learning Engineers to join our team. You are someone who loves solving complex problems, enjoys the challenges of working with huge data sets, and has a knack for turning theoretical concepts into practical, scalable solutions. You are a strong team player but also thrive in autonomous environments where your ideas can make a significant impact. You love utilizing machine learning techniques to push the boundaries of what is possible within the realm of Natural Language Processing, Information Retrieval and related Machine Learning technologies. Most importantly, you are excited to be part of a mission-oriented high-growth startup that can create a lasting impact. You will: Conceptualize, develop, and deploy machine learning models that underpin our NLP, retrieval, ranking, reasoning, dialog and code-generation systems. Implement advanced machine learning algorithms, such as Transformer-based models, reinforcement learning, ensemble learning, and agent-based systems to continually improve the performance of our AI systems. Lead the processing and analysis of large, complex datasets (structured, semi-structured, and
The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: We connect Airbnb’s community with the right information, in the right place, at the right time. We tailor Messaging & Notifications so hosts on Airbnb can streamline their operations, and travelers get just the information they need to enjoy their stay worry-free. Additionally, we are building new connections within our community to help enrich the experience of hosting & traveling on Airbnb: easing the process of hosting, and adding meaning to our guest’s trips. The data team utilizes industry-leading tools, builds scalable data systems and applies cutting-edge ML models to provide insights and empower all products in the Communication and Connectivity (CnC) organization. The Difference You Will Make: At CnC, data is foundational to our organization’s success.This role will lead key initiatives to design and build large-scale, distributed data systems - both batch and real-time processing. The data will power machine learning models and unlock new product features. You’ll be at the center of cross-functional collaboration, bridging backend, frontend/client, and machine learning engineering teams. CnC is applying GenAI and large language models (LLMs) to power products that enhance the Airbnb experience in various surfaces including highly used ones like Messaging. We're building a robust ML platform to power our product ambitions. A Typical Day: Shape the team’s long-term vision and roadmap in close collaboration with cross-functional partners across Airbnb Build strong relationships with partner engineering teams, including backend, client, data science, analytics,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: As a Software Engineer at Baseten, you will own one of the most critical surfaces of our business: pricing, billing, and revenue infrastructure. As we launch more and more products— billing is no longer just operational plumbing. It is a strategic lever for growth. This role will establish clear ownership of billing as a function and create leverage for Finance, Sales, and GTM teams while maintaining a seamless customer experience. RESPONSIBILITIES: Own Baseten’s end-to-end billing and revenue infrastructure, including pricing, invoicing, metering, and reporting foundations. Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams. Partner closely with Finance, Sales, GTM, and Forward Deployed Engineering to turn real-world workflows into reliable internal tooling and automation (quoting, approvals, renewals, usage reconciliation, revenue reporting). Design systems that scale with new products, packaging, and go-to-market motions, making billing a strategic lever for growth. Drive reliability and operational excellence for revenue-critical workflows: monitoring, alerting, incident response, backfills, and clear runbooks. Lead from the front on high-impact projects: clarify requirements, propose crisp technical approaches, ship iteratively, and raise the bar on quality and velocity. Debug and resolve
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role As a Principal Software Engineer, you will help define and build the high-performance, highly reliable foundation of the Nuro Driver. This role spans Device Platform, Performance, and Onboard Systems, requiring deep technical leadership across sensor and compute integration, onboard runtime systems, distributed execution, and autonomy software performance. You will architect hardware-agnostic device interfaces, inter-device communication pipelines, runtime APIs, and onboard software platforms that enable Nuro’s autonomy stack to run safely and efficiently across current and future vehicle platforms. You will also lead system-wide performance and reliability efforts, including latency reduction, resource efficiency, observability, automated validation, and tooling for debugging complex onboard systems. This is a highly cross-functional and high-impact role. You will partner with autonomy, hardware, safety, validation, opera
About the Role: As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams. Responsibilities: Design and build scalable, high throughput, and low latency distributed systems using Scala Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. Your Background: Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM b
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Who We Are: Shape the future of Roblox’s virtual economy. The Economy ML team is building the machine learning backbone that powers Roblox’s Marketplace, Developer Monetization, and Payments ecosystems. From intelligent pricing and personalized storefronts to dynamic layout optimization and avatar understanding, we’re reimagining how the Roblox economy drives user engagement, monetization, and creator success at scale. As a Principal Software Engineer (Data Systems) , you will architect, build and deploy high-scale, reliable real-time and batch data systems for personalization, search and recommendation across various product surfaces in Marketplace, Developer Monetization and Payments. You will be involved in key data projects from architecting event taxonomies and logging interfaces to real-time feature serving across multiple search and recommendation surfaces. What You’ll Do Act as data engineering lead for Economy ML, setting standards for batch vs streaming feature pipelines, table design, observability, and documentation used across the Economy group. Work as a hands-on contributor on our data systems to power content recommendation, search and personalization across Economy product
Role Overview Build reliable software services that power products, platforms, and business decisions. As a Senior Software Developer, you’ll design and deliver scalable applications, backend services, and integrations that perform well in production and evolve with changing business needs. You’ll apply strong software engineering practices across APIs, data-intensive applications, cloud services, AI-enabled solutions, and deployment pipelines. You’ll help shape technical solutions, improve system reliability, and contribute to a high-quality engineering culture. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design and develop scalable backend services and applications using Python or TypeScript. Lead the development of APIs, integrations, reusable software components, and AI-enabled features. Build reliable solutions for data ingestion, manipulation, service-to-service communication, and intelligent automation. Apply AI technologies and modern software engineering practices to improve product capabilities, developer productivity, and operational efficiency. Make sound technical decisions around architecture, performance, security, scalability, and maintainability. Deploy and operate applications using AWS services and CI/CD practices while improving testing, monitoring, documentation, and delivery standards. These are the essentials you’ll need to get an interview 5+ years of professional experience developing and delivering production software. Strong hands-on experience with Python; TypeScript or similar languages is also valuable. Proven experience building backend services, APIs, integrations, and service-oriented applications. Experience applying AI technologies, such as generative AI, machine learning services, intelligent automation, or AI-enabled application features. Strong understanding of software design principles, testing, debugging, performance optimization, and secure development. Experience working with cloud platfor
The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Senior Software Engineer, you will lead the development of a vertical feature or a horizontal capability to include defining requirements with stakeholders and implementation until it is accepted by the stakeholders. You will: Lead the design and implementation of scalable backend systems and distributed architectures for Federal customers. Manage the full lifecycle of feature development from requirement definition to deployment on classified networks. Direct the orchestration of asynchronous agent fleets to meet mission requirements. Lead customer engagements to translate mission needs into technical requirements. Own the communication with stakeholders to ensure implementation meets defined acceptance criteria. Conduct technical reviews and identify risks within machine learning infrastructure and model serving. Drive the platform roadmap by providing technical specifications for Federal product offerings. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in developing and deploying applications in a cloud-native environment. Understanding of containerization (e.g., Docker) and contai
Get new lead software engineer machine learning jobs by email
Daily job updates · Unsubscribe anytime