Jobs in United States

Ml Platform Engineer in United States

260 active opportunities · Updated October 2026

Explore current ml platform engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

C$100K – C$500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our IP delivery timelines are set as much by flow maturity as by design work. This role develops, deploys, and owns the RTL-to-GDSII methodology the IP physical design team runs on, so a new block, node, or customer variant starts from a working flow instead of a cold start. This role is hybrid, based out of Toronto, ON; Austin, TX, or Belgrade, Serbia. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A physical design or CAD methodology engineer who has built flows that production teams depend on daily. Automation-minded, happiest when you are removing manual steps and making PPA exploration repeatable. Building with AI as part of how you develop flows, and opinionated about where LLMs and ML-driven optimization genuinely help versus where they do not. An effective partner to design teams and EDA vendors, and a clear writer who documents flows well enough that others can run them without you. What We Need An Engineer with 5+ years developing and supporting physical design methodology or CAD flows in production use. Expertise with industry-standard tools (FusionCompiler/ICC2, Innovus/Genus, PrimeTime, RedHawk) and scripting languages (T

PythonAWSAISEM
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn from deployment and broadly distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. We aim to make our innovative tools globally accessible, transcending geographic, economic, or platform barriers. Our commitment is to facilitate the use of AI to enhance lives, fostered by rigorous insights into how people use our products. About the Role We are seeking Software Engineers (Emerging Talent) to join our Applied Engineering team. You’ll work in a highly iterative, collaborative, fast-paced environment to bring our technology to millions of users around the world, and ensure it’s delivered with safety and reliability in mind. We value engineers who are self-starters, care deeply about the end user experience, and take pride in building products to solve customer needs. In this role, you will: Own the development of new customer-facing ChatGPT and OpenAI API features and product experiences end-to-end Talk to users to understand their problems and design solutions to address them Collaborate with a cross-functional team of engineers, researchers, product managers, designers, and operations folks to create cutting-edge products Optimize applications for speed and scale Create a diverse and inclusive culture that makes all feel welcome. Your background looks something like: Bachelor's or Master’s degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 0-1 years of experience in software engineering or a relevant field Proficiency with JavaScript, React, and some backend languages (we use Python) Some experience with relational databases like Postgres/MySQL Interest in AI/ML (direct experience not required) Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadl

JavaScriptPythonJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Our mission at OpenAI is to discover and enact the path to safe, beneficial AGI. To do this, we believe that many technical breakthroughs are needed in generative modeling, reinforcement learning, large-scale optimization, active learning, and other areas. The team builds the performance-critical systems that allow OpenAI's models to run efficiently across a diverse set of AI accelerators. We work across the inference stack, from low-level kernels and compilers through model execution, to unlock the full capabilities of the underlying hardware. About the Role As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack required to run cutting-edge frontier models efficiently on the platform. This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution. You will work on the systems needed to support OpenAI's inference stack on Trainium, including developing and optimizing high-performance kernels, improving compiler support, and enabling efficient execution of the model forward pass. You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take full advantage of Trainium. The work may range from low-level hardware-specific optimization to compiler and runtime improvements to integrating new model architectures into the inference stack. If you enjoy working at the intersection of ML systems, compilers, kernels, and accelerator hardware, this role is for you. We're looking for engineers who are self-directed, comfortable operating across abstraction layers, and excited to solve challenging performance problems for frontier-scale AI systems. In This Role, You Will Build and optimize OpenAI's inference stack for AWS Trainium. Develop high-performance kernels for critical model operations and workloads. Extend and improve compiler support to efficiently target

AWSRestAIRust
M
📍 New York, new york, United States· Full-time
✓ Quality checkedCompany trend -67.9%

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed Engineers on our engineering team who want to work at the intersection of deep infrastructure work and direct customer impact. As an FDE, you'll partner with leading AI companies and foundation labs on cloud architecture, networking, storage, containerization, sandboxing, and more — helping them design and ship production infrastructure on Modal's platform. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the infrastructure stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and deploy massive-scale production workloads on Modal Lead technical discovery and architect

AWSAzureGCPDocker
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Making data driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide tooling and guidance to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. Engineers on Data Infrastructure are domain experts in Data Warehouse, Data Lakehouse, Spark, Workflow Orchestration, and Streaming technologies. We scale our existing data pipelines in a performant and cost efficient way while creating the necessary abstractions to make developing on top of this platform extremely simple for other engineers at Plaid. Responsibilities Contribute towards the long-term technical roadmap for data-driven and machine learning iteration at Plaid Leading key data infrastructure projects such as improving ML development golden paths, implementing offline streaming solutions for data freshness, building net new ETL pipeline infrastructure, and evolving data warehouse or data lakehouse capabilities. Working with stakeholders in other teams and functions to define technical roadmaps for key backe

PythonAWSMachine LearningAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The mission of the search and discovery team is to connect a billion users with the best connections and content at the right time, covering an extremely diverse range of recommendations and search problems at high scale. This is fundamental to how we personalize the experience to drive engagement and retention for Roblox users. Our recommendation systems suggest all games that the users see and play with across all surfaces. As a Senior Engineering Manager for the Search & Discovery ML team, you will lead the architects of exploration for one of the world's largest immersive platforms. You will be responsible for the core algorithms that connect over 120 million daily active users with millions of 3D experiences, items, and social connections. This is a high-impact leadership role where your team’s models directly determine the growth, retention, and satisfaction of the Roblox community. You will move beyond traditional recommendation systems into the realm of multimodal, agentic, and generative discovery . You Will: Lead and grow a high-performing team of machine learning engineers, fostering a culture of ownership, collaboration, and engineering excellence. Establish a lo

AWSGitMachine LearningAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Content Safety? As a Principal Machine Learning Engineer for Content Safety, you will define the future of proactive moderation, driving immense social impact through cutting-edge, innovative ML solutions, focused on critical and ambiguous safety challenges. You will set the 3-5 year technical strategy and architectural blueprint for how Roblox uses machine learning for content moderation. You will own the architectural and execution roadmap of massive-scale ML systems that mitigate violative UGC content before it impacts our community. You will feel a deep sense of responsibility in proactively protecting our community thoughtfully and fairly, while balancing user freedom with platform civility. Your efforts will ensure Roblox remains one of the safest places on the internet for our broad community of over 100 million daily active users. You will: Define and Own the Technical Vision: Define and lead the multi-year technical vision, architectural strategy, and execution for machine learning solutions in Content Safety, ensuring these systems proactively and effectively detect and mitigate violative content at massive scale. Strategic Stakeholder Partnership: Collaborate with execu

AWSGitMachine LearningAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Machine Learning Engineer within the Creator Services Machine Intelligence team, you will focus on the research and development of Embodied AI and Behavioral Agents that revolutionize how games are created and played on Roblox. You will bridge the gap between cutting-edge research and massive-scale product application, building agents capable of complex 3D gameplay and unblocking many use cases across Roblox, from automated playtesting to ensure quality, to "ML Players" with human-like movement and strategic reasoning, playing with real players in games. You will work on feature extraction, model training, building validation / RL platform as well as inference set up leveraging methods from imitation learning to reinforcement learning. And you will create generalizable agents that can perceive 3D environments, understand game rules, plan long-term strategies, and execute complex physics-based actions in real-time. You Will: Design and implement foundation models end to end through the feature extraction to inference for embodied agents. Define the long-term roadmap for Game AI and Embodied Intelligence, acting as a technical bar-raiser for code quality and architectural desig

PythonAWSGitMachine Learning
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $397.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. At Roblox , we’re building the tools and platform that empower a global community of creators and developers to build immersive experiences and a dynamic virtual economy. Our Economy ML team sits at the heart of this mission, delivering scalable machine learning systems that power personalization, pricing, search, and content understanding across all Economy surfaces: Marketplace, Developer Monetization, Payments, and Avatar. We’re looking for a Distinguished Engineer/Technical Director to lead the strategy and technical direction for ML systems , with a focus on large-scale recommendations, infrastructure, and emerging Generative AI applications. You’ll help build the systems that support retrieval, ranking, generative modeling, and LLM-powered personalization, all at massive scale. This role requires deep systems thinking, hands-on ML expertise, and a vision for how traditional ML and GenAI come together to power the future of the Roblox economy. Why Roblox for ML Systems AI/ML is a top company priority , with long-term investment. Real-world scale : Power millions of daily economic interactions across ranking, pricing, fraud, and search. Full-system ownership : Build and optimize end-to-

AWSGitMachine LearningAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $399.4K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Distinguished Machine Learning Engineer/Technical Director in the Safety organization at Roblox, you will drive the overall technical vision and execution for all machine learning initiatives focused on maintaining the safety and civility of our users. Our industry-leading safety features ensure Roblox remains a safe and inclusive environment for our community to express themselves creatively and share experiences without fear. The Safety org is the reason why Roblox is the safest place on the internet, protecting users You will provide technical leadership on AI/ML efforts for Trust and Safety as the Roblox platform scales to serve different age groups and geographic locations. You Will: Own the technical direction and implementation of machine learning solutions for safety-related systems Lead and mentor other engineers, fostering a culture of technical excellence and inclusivity Break down long-term product requirements into iterative deliverable stages, ensuring continuous improvement Craft and build large-scale machine learning models with billions of parameters, ensuring production-readiness Facilitate challenging technical decisions across multiple teams, demonstrating empathy a

AWSGitMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea

AWSAzureKubernetesCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role On the Accelerators team, you will help OpenAI evaluate and bring up new compute platforms that can support large-scale AI training and inference. Your work will range from prototyping system software on new accelerators to enabling performance optimizations across our AI workloads. You’ll work across the stack, collaborating with both hardware and software aspects - working on kernels, sharding strategies, scaling across distributed systems, and performance modeling. You'll help adapt OpenAI's software stack to non-traditional hardware and drive efficiency improvements in core AI workloads. This is not a compiler-focused role, rather bridging ML algorithms with system performance - especially at scale. In this role, you will: Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms. Optimize large-scale model performance (LLMs, recommender systems, distributed AI workloads) for diverse hardware environments. Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators. Collaborate on optimizations at the model code level (e.g. PyTorch) and below to enhance performance on non-traditional hardware. Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization. Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures. Contribute to runtime improvements, compute/communication over

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: GTM Innovation is a product engineering team with a charter to automate 100% of digital knowledge work in OpenAI's GTM, so sellers spend more time directly with customers. AGI-level reasoning doesn’t mean organization-level transformation “just works” out of the box; orgs must be redesigned around abundant intelligence and persistent virtual coworkers. Our team builds and scales a fleet of virtual coworkers that operate as full-time members of the account team, and redefines how our human-first revenue organization interacts with their agentic teammates. About the Role We’re looking for product mindset software engineers to join the GTM Innovation team. As a product engineer on this team, you’ll help OpenAI meet the world at scale. You’ll partner closely with go-to-market teams to understand their workflows, identify leverage points, and ship novel solutions using OpenAI’s API platform. You’ll move quickly from prototype to production, and your work will directly shape how customers experience our technology in the field. This role is ideal for engineers who want to be close to users, own end-to-end outcomes, and help define entirely new categories of enterprise software. In this role, you will: Build high-impact applications and tools that accelerate OpenAI’s go-to-market efforts Work across the full product lifecycle for GTM: prototype, iterate, ship, and maintain Embed with Sales, Technical Success, and Revenue Operations to identify user needs and build for them Apply OpenAI’s models in novel ways to solve real-world customer and internal workflow problems Translate learnings into feedback for Applied and Research teams to inform product development You’ll thrive in this role if you: Have 4+ years of experience as a software/ML/product engineer working on user-facing systems Former founder, or early engineer at a startup who built a product from scratch is a plus Are fluent in Python or JavaScript and comfortable building full-stack applications

JavaScriptPythonJavaAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: GTM Innovation is a product engineering team with a charter to automate 100% of digital knowledge work in OpenAI's GTM, so sellers spend more time directly with customers. AGI-level reasoning doesn’t mean organization-level transformation “just works” out of the box; orgs must be redesigned around abundant intelligence and persistent virtual coworkers. Our team builds and scales a fleet of virtual coworkers that operate as full-time members of the account team, and redefines how our human-first revenue organization interacts with their agentic teammates. About the Role We’re looking for Full Stack software engineers with a product mindset to join the GTM Innovation team. As a product engineer on this team, you’ll help OpenAI meet the world at scale. You’ll partner closely with go-to-market teams to understand their workflows, identify leverage points, and ship novel solutions using OpenAI’s API platform. You’ll move quickly from prototype to production, and your work will directly shape how customers experience our technology in the field. This role is ideal for engineers who want to be close to users, own end-to-end outcomes, and help define entirely new categories of enterprise software. In this role, you will: Build high-impact applications and tools that accelerate OpenAI’s go-to-market efforts Work across the full product lifecycle for GTM: prototype, iterate, ship, and maintain Embed with Sales, Technical Success, and Revenue Operations to identify user needs and build for them Apply OpenAI’s models in novel ways to solve real-world customer and internal workflow problems Translate learnings into feedback for Applied and Research teams to inform product development You’ll thrive in this role if you: Have 4+ years of experience as a software/ML/product engineer working on user-facing systems Former founder, or early engineer at a startup who built a product from scratch is a plus Are fluent in Python or JavaScript and comfortable building full-s

JavaScriptPythonJavaAWS
🔔

Get new ml platform engineer jobs in United States by email

Daily job updates · Unsubscribe anytime