Jobiba hiring network

Ml Platform Engineer Jobs

832 active opportunities · Updated for October 2026

Fresh results

14 shown

Explore current ml platform engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenTeams
📍 United States - Remote• Remote• $85K – $120K/yr
9 days ago

Who We Are Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it. OpenTeams exists to make ownership possible. Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves. If that sounds like your kind of work, we'd like to meet you. Senior Infrastructure Engineer - AI/ML Platform Location: U.S - Remote Work Authorization: U.S. citizenship required Clearance: An active clearance is not required at the time of hire. Candidates must be able to obtain and maintain a Secret security clearance, which includes a federal background investigation. Salary Range: $85,000 - $120,000 USD, dependent on experience level and location About the Role We're looking for a Systems Engineer to join our platform support team. This is a role for someone early in their infrastructure career who wants real ownership of production systems and is willing to earn it by being the person others can count on. The work runs across two streams. Internal Platforms covers the systems OpenTeams itself runs on — identity, cloud, internal applications, our website. You own that surface end to end, which means you're the first responder when something breaks and the one who fixes the underlying cause so it stops happening. Project Support means backing up our infrastructure engineers on the platforms we build and operate for customers, including government programs. Both streams start at the front line. That's the entry point, not the destination - as you build real context, the scope grows toward environment builds, deployment configuration, and eventually owning pieces of a platform outright. The two sides

REMOTEpythonsqlaws
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin

pythondockermachine learning
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Are you passionate about advancing the application of artificial intelligence? We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. Apply and scale optimization techniques across a wide range of ML models, particularly large language models. Collaborate with a diverse team to design and implement innovative solutions. Own projects from idea to production. REQUIREMENTS Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. Experience with one

pythondockerkubernetes
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Forward Deployed Engineer at Baseten, you will partner directly with customers to architect, build, and deploy high-scale production AI applications on Baseten’s platform. You’ll own the journey with customers from initial exploration to production deployment, translating ambiguous business goals into reliable, observable services with clear quality, latency, and cost outcomes. This role is a great fit for entrepreneurial engineers who want a front-row view into how modern companies adopt AI at scale and who enjoy working across product, software development, performance engineering, and customer-facing implementations. To be clear, this is an engineering role with hands-on coding and software development that also includes aspects of product management, technical customer success, and pre-sales solution engineering mixed in. EXAMPLE INITIATIVES Take a look at these blog posts written by members of our Forward Deployed Engineering team: Forward Deployed Engineering on the frontier of AI The fastest, most accurate Whisper transcription Deploy production-ready model servers from Docker images Deploy custom ComfyUI workflows as APIs RESPONSIBILITIES Develop and maintain software systems and product features using one or more general-purpose programming languages in a production-level environment, with a preference for Python due to its relevance in ML projects. Drive customer impact by designing, implementin

pythondockermachine learning
View job →
R
Roblox
📍 San Mateo• Full-time• From $295.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As a Model Optimization engineer on ML Platform, you will be responsible for digging deep into model internals to optimize performance, for both training and inference. We are looking for accomplished engineers to help us maximize performance of our platform. You Will: Optimize machine learning models for performance on GPU architectures, focusing on both training and inference workflows. Conduct low-level performance profiling analysis to identify bottlenecks in existing machine learning pipelines and propose actionable improvements. Contribute to the development of best practices and tooling for model optimization and deployment. Collaborate with cross-functional teams, including data scientists and software engineers, to integrate and deploy optimized models into production environments. Partner across organizations to build tooling, interfaces, and visualizations that make the ML@Roblox a delight to use. You Have: 6+ years of professional experience and a tool chest of system design experience upon which to draw to build performant system

awsgitmachine learning
View job →

The Opportunity Typography is central to how ideas are communicated. If you're passionate about beautifully created design, have deep curiosity about what AI can do, and take personal responsibility for creating products that people love; then this may be the role for you. Adobe Fonts supports millions of creatives in choosing and using typefaces across fonts.adobe.com, Express, Photoshop, Illustrator, Acrobat, and more Creative Cloud platforms. Our Internal Services team provides the platform engineering and deployment backbone for all of these. We manage CI/CD, deployment approaches, the services and data layers our engineers depend on, our observability and security stance, and increasingly the agentic tools that transform how our entire organization delivers software. We're seeking a Senior Software Development Engineer to lead this exciting journey in our San Francisco location. What you'll do Own and evolve our deployment platform. Lead strategy for CI/CD, PR environments, and release safety across a mixed fleet that includes containerized services, serverless services, and static front ends. Build the foundation for AI-accelerated development. Help build our agent factory and grow our internal agentic toolkit and skill library. Ship inference applications at scale. Take greenfield services from spec to production and standardize our ML/inference footprint. Modernize our services for the AI era. Identify where an existing service is held back by its current build and lead the fix. Rethink our security posture for agentic threats. Lead how we secure autonomous agents and their tool use. Expose Adobe Fonts to the agentic ecosystem. Extend our Model Context Protocol (MCP) surface and conversational, intent-based font discovery. Work higher up the stack, too. Contribute directly to search, browse, discovery, and the customer-facing experie

javascripttypescriptpython
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
A
Airbnb
📍 United States• Full-time• From $204K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: We connect Airbnb’s community with the right information, in the right place, at the right time. We tailor Messaging & Notifications so hosts on Airbnb can streamline their operations, and travelers get just the information they need to enjoy their stay worry-free. Additionally, we are building new connections within our community to help enrich the experience of hosting & traveling on Airbnb: easing the process of hosting, and adding meaning to our guest’s trips. The data team utilizes industry-leading tools, builds scalable data systems and applies cutting-edge ML models to provide insights and empower all products in the Communication and Connectivity (CnC) organization. The Difference You Will Make: At CnC, data is foundational to our organization’s success.This role will lead key initiatives to design and build large-scale, distributed data systems - both batch and real-time processing. The data will power machine learning models and unlock new product features. You’ll be at the center of cross-functional collaboration, bridging backend, frontend/client, and machine learning engineering teams. CnC is applying GenAI and large language models (LLMs) to power products that enhance the Airbnb experience in various surfaces including highly used ones like Messaging. We're building a robust ML platform to power our product ambitions. A Typical Day: Shape the team’s long-term vision and roadmap in close collaboration with cross-functional partners across Airbnb Build strong relationships with partner engineering teams, including backend, client, data science, analytics,

machine learningai
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications. You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Model Performance team: Baseten Embeddings Inference: The fastest embeddings solution available The Baseten Inference Stack Driving model performance optimization RESPONSIBILITIES Core Engineering Responsibilities Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap Performance & Innovation Impl

awsmachine learningai
View job →
R
Roblox
📍 San Mateo• Full-time• From $153.1K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Software Engineer on the Communications team, you will work on real-time communication across 2D and 3D spaces for billions of users while keeping interactions civil, safe, and expressive. Communication is at the core of what makes Roblox a vibrant social platform, the systems you build will define how tens of millions of people connect every day. You will work with teams across Game Engine, Safety, and ML Platform, building backend services at scale, data pipelines, and ML models that strengthen the underlying chat and policy infrastructure. You will build and ship detection systems that identify chat violations and continuously evolve to outpace bad actors circumventing our safety policies. As a core contributor to a broader suite of communication products, you will drive feature development while embedding safety into the foundation of the platform. If you are a developer with an understanding of large scale systems and want to shape how the next generation of Roblox users interact safely in the metaverse, you’ll be right at home within our highly-skilled and rapidly growing team. You will: Engineer and scale the communication infrastructure that supports millions of c

sqlawsgit
View job →
B
15 days ago

At Breeze, we're building the AI-powered infrastructure layer for global commerce, making it radically simpler for businesses to sell, get paid, and operate across markets. We go far beyond traditional payment processing. Breeze combines global payments, AI, stablecoins, and a Merchant of Record-like model to take on the complexity businesses typically manage themselves, including compliance, risk, fraud, chargebacks, reconciliation, and customer support. Our goal is simple: let businesses focus on building and selling great products while Breeze handles the complexity behind getting paid. Backed by Sequoia Capital , Multicoin Capital , and The Chainsmokers , Breeze is a successful, rapidly growing, and exceptionally well-capitalized company. We have the runway to think long term while remaining early enough that every person joining today can have a meaningful impact on what we build. We are hiring a Staff Machine Learning Engineer, Risk! As our Staff Machine Learning Engineer, Risk, you'll lead the evolution of our ML platform for payment risk, building the production-grade capabilities behind feature engineering, model training, deployment, monitoring, and continuous improvement. Risk decisions sit at the center of our business, and you'll own how those models get built, shipped, and kept healthy. This role reports to the CTO. You'll work closely with Risk, Software Engineering, and Data Engineering, and you'll be the senior technical voice for ML on the risk team. We're looking for someone who thrives in fast-moving environments, wants meaningful ownership, and is excited to build rather than simply maintain. What You'll Do Design and build ML infrastructure for payment risk detection, using Databricks as the core platform, in close partnership with software and data engineers. Bring structure to the team's ML environment: feature pipelines, versioning, job orchestration, and monitoring. Design and productionize models rather than just prototype them, including

machine learningaigo
View job →
R
Reddit
📍 Ontario• Full-time• Remote
15 days ago

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com . Reddit has a flexible workforce! If you happen to live close to one of our physical office locations our doors are open for you to come into the office as often as you'd like. Don't live near one of our offices? No worries: You can apply to work remotely in any country in which we have a physical presence. Team Description Reddit is poised to rapidly innovate and grow like no other time in its history. We’re currently hiring across multiple teams, some of these teams include: Ads ML Serving Team The Ads ML Serving team is part of Reddit’s Ads ML Platform, which builds the infrastructure and tools that power machine learning across Ads. This team focuses on creating a highly reliable, scalable, and efficient ML serving stack. Their work includes evolving long-term serving architecture, integrating closely with the ads serving stack, optimizing CPU/GPU performance, and building model velocity tools like observability libraries and model quality gating. Attribution & Identity Team The Attribution & Identity team builds products that help advertisers understand and measure the impact of their campaigns. They focus on attribution systems, identity solutions, and advertiser experimentation tools that improve performance insights and usability. Their goal is to make Reddit’s advertising platform more effective, transparent, and data-driven. Ads Growth Team The Ads Growth team drives initiatives to expand Reddit’s advertiser base, with a focus on Small to Medium Businesses (SMBs). We build and scale the technical founda

REMOTEpythonjavaredis
View job →
G
Godaddy
📍 Ontario• Full-time• From C$169K/yr
1mo ago

Location Details: BC or ON - Canada, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team Join GoDaddy, where your work will directly influence how our engineers craft and build, and how millions of entrepreneurs utilise our products. Our role sits at the intersection of AI, platform engineering, and user-centred design, enabling a rare opportunity to improve both internal developer efficiency and the customer experience. If you have a proven track record of building high-leverage platforms, delivering exceptional user experiences, and can push the boundaries in advancing the possibilities of AI-assisted efficiency, we encourage you to apply. What you'll get to do... Own GoDaddy’s developer experience strategy, driving AI efficiency, automation, and reduced cycle time Champion modern, agentic workflows, using copilots, autonomous agents, intelligent CI/CD, and automated testing Build the evolution of developer platforms including local development, CI/CD, testing frameworks, and observability. Identify and eliminate friction across the end-to-end engineering lifecycle Establish metrics and dashboards for engineering velocity, efficiency, and developer fulfilment Drive and own our enterprise UX Platform vision, including design systems, component libraries, pattern governance, and accessibility standards Partner with Product, UX, and AI/ML to enable AI-powered and conversational UX patterns across GoDaddy product surfaces Ensure cross-product experience consistency across web, mobile, and emerging surfa

ci/cdaigo
View job →
N
Numa
📍 Montreal• C$200K – C$250K/yr
10 days ago

Senior Machine Learning Developer Location: Montreal, Quebec, Toronto, or Ontario About Numa Numa is building the platform to power AI-native dealerships, rearchitecting automotive service and sales with advanced AI agents that automate customer interactions, streamline operations, and reimagine how dealerships work. Numa integrates AI into every aspect of dealership functions—from rescuing customer calls and voicemails that generate more revenue, to reducing customer resolution times that drive overall customer satisfaction (CSI), to improving dealership team productivity and accountability. Numa has raised $50 million from leading investors (Google, Threshold, Costanoa, Mitsui, and Touring Capital). The Role We’re hiring a Senior Machine Learning Developer to build and ship ML/AI systems that interact with real customers thousands of times a day. Our voice agents book service appointments, rescue missed calls, and route callers through natural conversations. You’ll work across products and platforms. You’ll ship AI features including prompts, agents, tools, and production ML models, while building the evaluations and tooling that help teams ship with confidence. You’ll also contribute to our ML platform, including model serving, LLM infrastructure, and production observability. At Numa, we believe great ML is about more than building bigger models. It’s about knowing whether a change is good enough to ship. Our evaluation-first approach makes that measurable in CI and production. What You’ll Do Build conversational AI systems for phone and SMS that understand customer needs, take action, and know when to act autonomously Develop tooling such as memory, knowledge graphs, and validated customization that help agents reason and adapt to dealership needs Train, evaluate, and deploy ML models using Ray Serve and Dagster for prediction, classification, ranking, and capacity forecasting, and keep them healthy in production Create offline and online evalua

pythongcpkubernetes
View job →
🔔

Get new ml platform engineer jobs by email

Daily job updates · Unsubscribe anytime