Jobs in United States

Ml Platform Engineer in United States

260 active opportunities · Updated October 2026

Explore current ml platform engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $234K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Staff Software Engineer with deep experience in GenAI/ML to join Datadog’s Application Performance Monitoring (APM) team. APM is a product which provides deep visibility into applications, enabling users to identify performance bottlenecks, troubleshoot issues, and optimize services. With distributed tracing, profiling, out-of-the-box dashboards, and seamless correlation with other telemetry data, Datadog APM provides some of the deepest and most structured visibility into the health and performance of applications. This context sets us up for an opportunity to be the world leaders in agentic investigations and incident troubleshooting. You’ll act as a technical leader within the APM group, focused on agentic workflows. You’ll lead efforts to design, train, evaluate, and deploy GenAI/ML models at scale. We’re looking for a product-minded ML engineer with strong technical expertise, excellent communication skills, and a track record of driving impactful initiatives end to end. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical leader within the APM organization, driving GenAI/machine learning projects from concept to production. Build and benchmark GenAI/ML models using state-of-the-art techniques. Collaborate with cross-functional teams to build automated investigation and triaging tools. Influence product direction by bringing a strong product mindset to your work, always advocating for the end user. Guide teams through ambiguity, scaling challenges, and evolving requirements with clear technical direction. Actively mentor engineers and influence engineering culture through leadership in design reviews, technical talks, and working groups. Who You Are: You have a BS/MS/PhD in a scientific field or equiva

Machine LearningAIGoRust
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services. What you will be doing: Leading the overall architecture and design of our distributed storage service optimized for AI/ML Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes. Develop and maintain user-space applications, containers, Go-bindings, and CLI tools. Building features for a distributed storage service to enhance availability and reliability for large-scale deployments Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services. Automating distributed storage service end-to-end, including deployment, management, and monitoring What we need to see: Bachelor’s of Science in Computer Science, or related field (or equivalent experience) with 8+ years of industry experience Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations Strong track record of delivering distributed services in a variety of distributed computing environments Experience in i

KubernetesArtificial IntelligenceAIGolang
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $212K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Our web and API surfaces handle requests from guests and hosts alongside a growing volume of automated agents: AI assistants, crawlers, and scrapers. We build the systems that bring clarity to this traffic, combining in-house ML and vendor signals to decide in real time how to serve billions of daily requests. Anti-bot and anti-scraping detection is our most adversarial mandate, but the wider challenge is full traffic classification: building evaluation frameworks that tell legitimate automation apart from abusive actors, so high-stakes decisions hold up across the fleet. The Difference You Will Make: You will architect and maintain Airbnb’s end-to-end traffic classification ML systems, balancing high-performance model deployment with rigorous offline data pipelines. Success is measured by your ability to harden edge-traffic policies—targeting reduced bot-incident MTTM—and by establishing rigorous evaluation practices that ensure foundational signal accuracy and evasion-resistance across the fleet. A Typical Day: Own the complete lifecycle of traffic-scoring models, from problem framing to real-time deployment, managing the adversarial feedback loop to ensure high evasion-resistance and directly drive reductions in bot-incident MTTM. Architect robust offline-to-online pipelines that produce certified source-of-truth datasets, establishing rigorous evaluation frameworks—such as stratified benchmarks and leakage-prevention checks—to ensure every model improvement is empirically measurable and defensible. Execute model optimization within strict millisecond latency budgets at the

SQLGitMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the growing set of safety challenges as AI becomes more powerful and used in more settings. Key focus areas include how to train nuanced safety behaviors, how to make the model robust to bad actors, how to address privacy and security risks, and how to make the model trustworthy in safety-critical situations. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. About the Role We’re seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You’ll advance safety post-training and robustness, helping models follow nuanced policies while preserving their usefulness and capabilities. In this role, you will: Research and implement methods for safety training, reinforcement learning, and adversarial robustness. Develop evaluations, identify model failure modes, and use findings to improve training. Work with research, engineering, security, and policy partners to support safe, reliable deployment. You might thrive in this role if you: Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness. Have a degree in computer science, machine learning, or a related field, and strong deep learning research or engineering skills. Have experience improving model safety for deployment and enjoy collaborative research. Are motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings. Security Requirements Active TS/SCI clearance or equivalent. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefi

AWSRestMachine LearningAI
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect

PythonAWSMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Proactivity Research team, within OpenAI’s broader Personal AGI team, is focused on making our models in ChatGPT and future potential products proactive in ways that are truly useful. We're laying the technical foundations for AI that can anticipate what users need in real time, adapt as their goals and preferences shift, and build a deeper, evolving understanding of the person it's helping. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models’ personalization and agentic capabilities. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a highly personalized, collaborative, and proactive assistant. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve the proactivity and ability of our models to further user goals. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the other research and product teams to influence the shape of technical solutions in the product You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of LLM post-training and evaluation approaches Are passionate about, or have experience thinking about, personalization and enabling users to achieve their goals Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI

AWSRestMachine LearningAI
P
📍 California, United States
✓ High-confidence listingCompany trend +671.4%

$176.6K – $294.3K/yr

Quick readStrong listing-quality and freshness signals

ROLE SUMMARY As an AI-enabled strategic scientific leader, you will partner with Oncology Research Leadership to integrate AI-driven strategies that will drive data-informed decision-making across our Oncology Research pipeline. As the AI Portfolio Lead, you will bridge cutting-edge AI/advanced analytics with deep Oncology R&D expertise to guide our portfolio strategy. Working as an individual contributor reporting into the Head, Portfolio Strategy and Program Management, you will analyze and continuously assess Oncology Research projects using AI-driven insights (predictive models, scenario simulations, competitive intelligence) and propose actionable strategies to Oncology Research Leadership. This role is science-focused and spans across discovery and preclinical stage programs, ensuring that our project teams pursue the most promising scientific approaches and mechanisms of action. Your work will directly inform pipeline prioritization and resource allocation decisions, enhancing the quality and objectivity of governance deliberations with robust data. ROLE RESPONSIBILITIES Portfolio Analysis and Insights: Continuously analyze the oncology Research portfolio (ESD through Preclinical) using advanced AI/ML tools and analytics to evaluate each program’s scientific strength, probability of success, and strategic portfolio fit per disease area strategies. Strategy Recommendation: Develop and propose data-driven portfolio strategies (at both project and portfolio levels) to senior leaders and Research governance committee. Use scenario analysis and predictive modeling to highlight optimal project prioritization, pipeline balance, and resource allocation scenarios. AI-Enabled Decision Support: Integrate AI-derived insights (e.g. machine learning predictions, knowl

Machine LearningAIRecruitment
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Principal ASIC Design Engineer About the Role At Micron, we transform how the world uses information to enrich life for all. As a global leader in memory and storage innovation, we develop technologies that accelerate intelligence and enable the next era of AI, ML, and advanced computing. Micron’s Interface Pathfinding team drives performance-scaling innovation across circuits, signaling, packaging, and interconnects with a 3–5 year technology horizon. A core part of that work is silicon-based validation of novel PHY solutions, and we are growing the team to execute. As the Principal ASIC Design Engineer , you will be the primary digital contributor on a deliberately small, senior team — united around the goal of carrying high-speed interface technologies from architecture to tape-out. The team’s analog and chip-level architecture is anchored by a deeply experienced analog custom design engineer; your role is to be the authoritative digital voice — owning RTL design and micro-architecture of the digital blocks, defining timing constraints, supporting verification, and serving as the key interface between the digital design and the analog and layout specialists who will carry the implementation through to silicon. On a test ASIC of this scope, the front-end digital work is the critical path, with contractor and layout support engaged as bandwidth demands warrant. What You’ll Own Digital Block Architecture & RTL Design Own the micro-architecture and RTL implement

AIRecruitment
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Principal Signal Integrity Engineer — Interface Pathfinding About the Role At Micron, we transform how the world uses information to enrich life for all. As a global leader in memory and storage innovation, we develop technologies that accelerate intelligence and enable the next era of AI, ML, and advanced computing. Micron’s Interface Pathfinding team operates at the leading edge of that mission — serving as a bridge between corporate strategy and engineering through forward-looking investigation and development of energy-efficient bandwidth solutions. Our work spans innovative circuit, signaling, packaging, and interconnect solutions with a 3–5 year technology horizon, grounded in rigorous analysis, hands-on measurement, and a commitment to building a robust intellectual property portfolio. As a Principal Signal Integrity Engineer , you will be a core technical contributor on a deliberately small, senior team united around the goal of preparing high-speed interface innovations for high-confidence product adoption. What You’ll Work On Micron’s Interface Pathfinding team investigates and validates novel interconnect and signaling solutions for high-speed memory and die-to-die interfaces. The signal integrity scope is broad and analytically deep — spanning electrical modeling, channel analysis, and substrate evaluation — with an emphasis on original analysis rather than execution of established playbooks. Key signal integrity domains include: Intercon

Artificial IntelligenceAIRecruitment
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. We are the Data team within Plaid’s Fraud organization. We build the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s network data to help identify and prevent fraud before it happens. Our team owns the end-to-end ML lifecycle, from feature pipelines and model training to production serving and monitoring, ensuring our systems are reliable, scalable, and built to support hundreds of customers and data partners. As a Data Scientist on the Fraud Data team, you will analyze customer and network traffic to understand how Plaid Protect performs across a range of use cases and customer segments. You’ll build dashboards and metrics that provide a clear, shared view of product performance, run backtests to evaluate performance and identify high-impact rules and model strategies, and generate insights that support customer growth and expansion. You’ll also design scalable data models and schemas to enable reliable analysis and reporting, while partnering closely with Product and Engineering to design and analyze experiments for new customer-facing features. Responsibilities: Work at the intersection of product analytics, machine learning, and fraud a

PythonSQLAWSMachine Learning
MT
📍 Richardson, TX, United States
✓ Quality checkedCompany trend +1266.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Sr Design Engineer, you will work on design, simulation, and validation of next ‑ generation High ‑ Bandwidth Memory (HBM) architectures and circuit blocks. HBM requires advanced DRAM design knowledge combined with deep understanding of 3D stacked architecture, TSV signaling, wide I/O interfaces, PHY timing, power integrity, and system co ‑ optimization with GPUs/accelerators. This role sits at the intersection of DRAM design and high ‑ performance computing, enabling future AI/ML, HPC, and advanced graphics products. In this position, you will collaborate with Micron’s various design and verification teams all over the world and support the efforts of groups such as Product Engineering, Test, Probe, Process Integration, Assembly and Marketing to proactively design products that optimize all manufacturing functions and assure the best cost, quality, reliability, time-to-market, and customer satisfaction.

AIRecruitment
M
📍 Richardson, TX, United States
✓ Quality checkedCompany trend -75%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Our vision is to transform how the world uses information to enrich life for all. Join an inclusive team passionate about driving innovation for next-generation memory solutions powering AI/ML and high-performance computing systems. As a Director of HBM RTL Design and Integration, you will lead a team responsible for the design, integration, and delivery of next-generation HBM SoC logic die, with a strong focus on RTL development and IP integration. You will drive technical strategy and execution across multiple product generations, working closely with architecture, verification, physical design, firmware, and product engineering teams. Key Responsibilities Lead SoC RTL design and integration for HBM logic die, including subsystem partitioning, IP integration, and SoC-level design convergence. Drive translation of architectural and micro-architectural specifications into robust, high-quality RTL implementations across multiple teams. Oversee SoC integration aspects, including clocking, reset, power intent, configuration infrastructure, and system-level design correctness. Establish design methodologies and best practices to improve quality, reuse, and development efficiency across HBM programs. Partner with SoC Architecture, Verification, Physical Design, Firmware, and System teams to ensure successful end-to-end product execution. Work closely with Product Engineering, Test, Probe, Process Integration, Assembly, and Manufacturing to ensure robust, manufacturable HBM

PythonAIRecruitment
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

As a Product Marketing Manager for NVIDIA NIM, you will play a pivotal role in helping enterprises deploy AI. NVIDIA NIM is revolutionizing AI from experimentation to production, packaging optimized, production-ready models for rapid deployment. Our product marketing team operates at the intersection of developer tooling and enterprise go-to-market, ensuring that our narrative resonates with ML engineers and CTOs alike. If you relish diving into technical depth and translating it into compelling stories, this is the perfect opportunity! What you will be doing: Launching products: You'll build and manage launch plans for high-visibility NIM releases, coordinating assets, timelines, keynote slides, demos, and press materials while driving cross-functional execution with product management, engineering, technical marketing, and PR or equivalent experience. You'll align the team on messaging and positioning for NIM across developer and enterprise audiences. You will develop messaging docs, sales enablement materials, customer presentations, solution overviews, and web content. Driving awareness: You'll identify target audiences and content gaps, then build the assets that fill them — blogs, webinars, demos, solution briefs, and more. Crafting the ecosystem story: You'll drive co-marketing engagements with NVIDIA's model providers, cloud partners, and ISV ecosystem to showcase the full range of possibilities with NIM. Collaborating with PR: You'll work with PR on press launches to ensure the NIM story is accurate, compelling, and consistent across every channel. What we need to see: Excellent written and verbal communication skills, with a proven track record of articulating technical value to both developer and executive audiences. College degree or equivalent experience. More t

S
📍 New York, United States· Full-time· Remote
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Summary The Senior Flex Solution Engineer (SE) is a high-impact, customer-facing technical role supporting AI/Machine Learning (AI/ML) initiatives across our strategic US Major Accounts. This role is designed for a technical professional who possesses a strong foundation in AI/ML concepts, technologies, and solution architecture, positioning them above a generalist but not necessarily requiring deep specialization (Level 200-300 technical depth). The Flex SE will act as a critical, hands-on technical resource, accelerating customer adoption and success. This role translates business challenges into AI/ML- driven solutions through the rapid development of MVPs and prototypes, supporting our regional SE teams, and driving growth in this strategic area. The AE and SE maintain ownership and ultimate approval over the account strategy. The Flex SE role is designed to be supportive, not to supersede their authority. The RVP should provide prescriptive guidance on which accounts the Flex SE should prioritize, as the RVP possesses the most comprehensive regional overview. Key Responsibilities and Scope Technical Leadership & Solutioning Partner with regional Account Executives (AEs) and Solution Engineers (SEs) to identify, qualify, and develop AI/ML opportunities within USMajo

Machine LearningAIGoRust
V
📍 United States· Full-time
✓ Quality checkedCompany trend -88.6%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Data Engineer, you’ll be responsible for laying the foundation for a best-in-class analytics function. You’ll partner closely with our engineering team and business stakeholders to ensure that our analytics stack and processes meet the business needs today with an eye towards the future. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as a Senior Data Engineer at Vanta: Design and deploy data infrastructure needed to drive data-driven decision-making solutions Design and implement complex data orchestration models, modeling metadata, scaling reporting tools for data science and ML products users Be the company’s expert on data administration, data management and scalable data systems Write highly tuned, scalable SQL queries running over large-scale, heterogeneous data warehouses Work with the Product and Enterprise Engineering system teams to structure source systems for reporting consumption across the enterprise Help maintain CDC pipelines to power customer reporting Help develop front end applications to expose analytical data sets enterprise wide How to be successful in this role: Have at least four years of experience working with data and two years of experience in Software Engineering or a related field. Have experience with common analytics tooling (e.g. Stitch/Fivetran, Snowflake/BigQuery/Redshift, dbt, Airflow, Dagster). Have good working knowledge of AWS data infra systems and Terraform. Bring a system-oriented and software engineering mindset to the Data Engineering practice. We’re looking to build frameworks that manage data, and minimize bespoke queries Deep kn

SQLAWSRestAI
🔔

Get new ml platform engineer jobs in United States by email

Daily job updates · Unsubscribe anytime