Jobiba hiring network

Ml Platform Engineer Jobs

832 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current ml platform engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Engineering Manager, AI Conversation Platform Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The newly formed Conversation Platform team aims to build a conversation platform for all merchants who use Stripe. We are doing so by (a) automating the easy tasks, and (b) assisting our users in the difficult tasks. Some examples include customizing the Stripe landing page to suggest bespoke integrations, allowing users to command the Stripe API in natural language, and resolving user issues automatically. We are developing RAG based systems on the latest LLMs as well as fine-tuning our own models. We’re an end-to-end team going from ideas to models to shipping in production. What you’ll do Responsibilities Driving an ambitious vision for AI/ML that benefits our users Setting the technical & process direction for the team based on business goals Brainstorm and coordinate product integrations with partner teams Proposing new ideas and building prototypes Be an integral part of a larger ML community internally & externally Hire & develop a world-class team to deliver high-quality ML systems. Coach engineers to help them grow in their careers and maintain a high bar Who you are We are looking for ML Engineering Managers who are passionate about using ML to improve products and delight customers. You have experience leading teams that develop streaming feature pipelines, build ML models, and deploy them to production, even if it involves making substantial changes to backend code. You a

machine learningaigo
View job →
T
Tenstorrent
📍 Austin• Full-time• C$100K – C$500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking an Physical Design Engineer to lead cross-functional efforts to solve complex physical design challenges and develop end-to-end RTL-to-GDS methodologies across advanced nodes, with a strong focus on PPA and runtime improvements. The engineer will architect, integrate, and deploy AI/ML-driven solutions into production physical design flows, creating custom CAD tools and partnering with internal teams and EDA vendors to drive next-generation, ML-enabled capabilities. This role is hybrid, based out of Santa Clara, CA or Austin, TX or Fort Collins, CO. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who you are BS in Electrical or Computer Engineering (or equivalent experience) with 5+ years in Physical Design CAD methodology at advanced nodes. Proven track record improving PPA and/or runtime on high-performance, low-power taped-out designs. Hands-on with industry-standard EDA tools (e.g., Fusion Compiler) across synthesis, P&R, STA, signoff, and hierarchical flows. Strong Python/Tcl and data skills, with interest or experience in ML frameworks (PyTorch, TensorFlow), and the ability to drive complex projects independent

pythonawsrest
View job →
DU
DoorDash USA
📍 San Francisco• Full-time
15 days ago

About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus

pythonjavasql
View job →
A
Airbnb
📍 Bangalore, India• Full-time
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. Senior Software Engineer(AI/ML),Trust - India The Community You Will Join: Everyone at Airbnb thinks about trust, but our team obsesses over it daily. At the core of trust is safety, and thus we spend a significant amount of our time and energy keeping the community safe. The Trust team is responsible for protecting our community and platform from fraud while also ensuring our hosts, guests, homes, and experiences meet our high standards. We constantly work to fight against online fraud (such as monetary loss, compromised accounts, spam and scam in messages, fake inventory, etc.) as well as offline fraud (theft, property damage, personal safety, etc.). We also work on onboarding and screening of users, and think about complex topics like identity and reputation to ensure that every interaction with Airbnb helps build trust in us and our community. Trust Engineering is responsible for the technology vision and development of a complex stack that runs on every key interaction on the platform. The Difference You Will Make: As part of the Trust Engineering team, you will work in a team of talented software engineers to help us build intuitive & delightful experiences, strengthen our current offerings, and deliver new products to strengthen our Trust defenses. You will play a significant role in

pythonjavakubernetes
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a

pythonawsrest
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches

REMOTEpythonawsazure
View job →
R
Roblox
📍 San Mateo• Full-time• From $260.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Safety? At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization’s mission is to become the leader in civil immersive online communities. We systematically and proactively work to detect, remove, and prevent problematic content and behavior. We seek to influence and shape the product roadmap and prioritization, build safety products, and measure our impact on the community of

awsgitmachine learning
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale is seeking a Staff Software Engineer to lead the technical vision for our Infrastructure team. As a Staff Engineer, you will be responsible for the architectural evolution of our control plane and data plane, ensuring that our "infinite laptop" vision scales to meet the most demanding distributed AI workloads in the world. You will act as a force multiplier, setting the standards for Kubernetes-based cloud-native infrastructure while mentoring engineers and driving cross-functional alignment across the Ray open-source community and our proprietary product teams. Key Responsibilities Architectural Leadership: Define and drive the multi-year technical roadmap for services that orchestrate Ray clusters across diverse cloud and on-premises environments. Systemic Optimization: Lead the design and optimization of high-performance control plane components specifically tailored for large-scale, heterogeneous AI/ML workloads. Platform Reliability: Establish the organization-wide standards for the reliability, scalability, and observability of Anyscale-managed infrastructure. Strategic Integration: Direct the long-term strategy for accelerator integration (GPUs, TPUs) and container management to ens

pythonawsazure
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We’re a team of high-output generalists where ML and systems engineering converge. This is not a "run the models" role. We reason from first principles about why a perception model learns what it learns, close the gaps that cap its performance, and raise the bar on the data and evaluation loop that drives autonomy. Your work will directly impact how autonomous systems understand rare scenarios, adapt to global geographies, and scale safely. About the work You’ll solve autonomy’s hardest data challenges through applied ML and systems rigor: Diagnose why perception models underperform on the long tail, and turn that into targeted data and training priorities. Design eval metrics and regression detection that tell us whether a model is ready. Curate and clean training data for segmentation and occupancy; hunt the data problems that silently cap performance. Run controlled ML experiments and ablations; cleanly sepa

pythonrestai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We are looking for a Senior/Staff Software Engineer to serve as a technical leader for Nuro’s ML Data engine. You will sit at the critical intersection of Autonomy, Machine Learning, and Infrastructure, acting as an architect for the systems that feed our autonomy AI models. In this role you will be a member of the Autonomy team responsible for executing the technical strategy for transforming massive amounts of autonomy data into high-value training signals for autonomy decision making. You will design and build data products for autonomy researchers, develop queries for rare "needle-in-a-haystack" scenarios, and trigger labeling and data ingestion workflows without human intervention. You will partner directly with Autonomy ML researchers to understand their data needs, collaborate with infrastructure teams to define the right data interfaces and APIs, and build robust data selection, simulation, and introspe

pythonmachine learningai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro is seeking a Software Engineer with expertise in large-scale infrastructure, workload orchestration, and data processing to join our ML Infrastructure team . In this role, you will focus on building and evolving the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management to accelerate the Nuro Driver™ development lifecycle. About the Work You will build the foundation that powers Nuro’s model development from experimentation to production. Key responsibilities include: Resource Provisioning & IaC: Scaling automated infrastructure-as-code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments. Intelligent Scheduling: Designing and optimizing workload orchestration to maximize

redisawsazure
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro is a self-driving technology company on a mission to make autonomy accessible to all. Founded in 2016, Nuro is building the world’s most scalable driver, combining cutting-edge AI with automotive-grade hardware. Nuro licenses its core technology, the Nuro Driver™, to support a wide range of applications, from robotaxis and commercial fleets to personally owned vehicles. With technology proven over years of self-driving deployments, Nuro gives the automakers and mobility platforms a clear path to AVs at commercial scale, empowering a safer, richer, and more connected future. About the Role The Mapping team in Nuro takes a machine learning-first path to unblock geographic capability with lower costs. This team plays a crucial role in the advancement of autonomous driving systems by creating and improving different components in the full lifecycle of ML models for multiple teams in this organization. We are searching for an engineer with experience building reliable and scalable machine learning infrastructure and a strong desire to contribute to the future of robot navigation for logistics and transportation. About the Work Mainly focus on building and improving HD map generation and release pipelines. Develop scalable workflows to manage first party and third party map data. Develop and manage APIs for internal users to access data. Improve and refactor existing workflows and toolings to boost efficiency. Engage with other mapping teams to help identify issues and establish long-term relationships that include knowledge sharing. About You BS in Computer Science, Robotics or another quantitative area. You have experience in one or more of the following areas: large-scale distributed systems; data storage and processing systems; advanced algorithms using C++ and Python; multithreading; and software performance tuning and optimization. Ability to efficiently develop, debug, and support new technologies in a changing environment. Str

pythonmachine learningai
View job →
R
Roblox
📍 San Mateo• Full-time• From $399.4K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Economy ML team sits at the center of this mission. We build ML and AI systems that power personalization, search, pricing, recommendations, and virtual item intelligence across Roblox Economy surfaces, including Avatar Marketplace, Payments, Developer Monetization and Subscriptions. As Roblox continues to invest deeply in AI and ML, this organization also plays a key role in advancing next-generation AI capabilities across the company through close partnership with Discovery, Foundation AI, Engine, Creator and product engineering teams. Our work spans both product impact and foundational ML innovation at Roblox scale. We are looking for a Director of Engineering to lead and scale Economy ML. This leader will define the strategy, organization, and execution model for a fast-growing ML organization spanning recommendation systems, ML infrastructure, search, content intelligence, and Generative AI applications. Why Roblox for ML/AI AI/ML is a top company priority , with strong executive support and long-term investment. Massive real-world scale: Millions of users, creators, and economic interactions every day. Unique technical challenges: Recommendations, pricing, fraud, search, and conte

awsgitmachine learning
View job →
S
1mo ago

Who We Are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies — from the world's largest enterprises to the most ambitious startups — use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the Team The Datalake team builds and maintains Stripe's foundational data access and governance infrastructure — the paved path for safe, fast, and compliant access to Stripe's critical big data assets. We serve developers, data engineers, analysts, ML and AI teams, security teams, and business users across the company. The team is in the middle of a significant architectural transition as Stripe grows. We are making Stripe's data lake a first-class citizen of the modern data ecosystem to support our growing scale and diverse workloads. What Makes This Role Compelling Foundational infrastructure with broad reach: The Datalake team's systems sit in the critical path of nearly every data workload at Stripe. Decisions affect petabytes of data, hundreds of production pipelines, and every engineering team that builds on Stripe's data lake. Active, high-stakes architectural transformation: The team is executing a multi-year migration to modern, OSS-aligned solutions — a technically deep project with real architectural choices at each step, including API design, compute engine integration, authorization model, and per-table credential vending. Active, high-stakes, OSS-aligned architectural transformation: You will lead a multi-year migration to modern, open-source solutions like the Apache Iceberg REST Catalog. This is a technically deep project involving critical architectural choices at each step, from API design and compute engine integration

azurerestai
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Cloud Support Engineer — AI/ML & Programmability (Night Shift) Location: Pune, India Snowflake's Support team is expanding. We're looking for a Cloud Support Engineer who enjoys working with data and solving a wide variety of problems, drawing on hands-on experience across operating systems, database technologies, big data, data integration, connectors, and networking. Our mission is to make Snowflake the preferred platform for running all AI, ML, data science, and data engineering workloads. You'll join a highly productive, fast-moving team supporting Snowflake Cortex and our ML product lines — work that is central to delivering on Snowflake's AI Data Cloud mission. Snowflake Support is committed to providing high-quality resolutions that help customers deliver data-driven business insights and results. We are a team of subject matter experts working collectively toward our customers' success, building partnerships by listening, learning, and connecting. Snowflake's values shape how we deliver world-class Support: putting customers first, acting with integrity, owning initiative and accountability, and getting it done. As a Cloud Support Engineer, you'll be the technical partner our customers turn to for guidance on using Snowflake effectively. You'll also be the voice

pythonkubernetesgit
View job →
🔔

Get new ml platform engineer jobs by email

Daily job updates · Unsubscribe anytime