Jobiba hiring network

Ml Platform Engineer Jobs

832 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current ml platform engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
Coinbase
📍 - USA• Full-time• Remote• From $253.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Staff Software Engineer on the Data Platform team within Platform , you'll define and lead the technical strategy for Coinbase's data infrastructure, spanning ingestion, transformation, warehousing, streaming, and serving systems. This is a foundational role at the intersection of distributed systems, data engineering, and AI-readiness, reporting to the Senior Director of Engineering. You'll set architectural direction, drive multi-quarter roadmaps, and transition the organization from managed-service dependency toward engineering-built, platform-grade infrastructure that powers everything from fraud detection to modern multi-agent AI architectures. What you'll do: Own the technical strategy and architecture for Data Platform, setting direction across data ingestion, transformation, warehousing, streaming, and serving systems while driving engineering-led cost reduction at the infrastructure layer. Architect data infrastructure to natively support AI and ML workloads, ensuring pipelines, data lake systems, and compute can power ML training, feature stores, real-time inference, and multi-agent AI architectures at scale. Drive the evolution to near-real-time data availability, enabling downstream teams across Coinbase to act on fresher data for fraud detection, financial reporting, and analytics. Build alignment and secure commitment from senior leadership

REMOTEawsaigo
View job →
O
1mo ago

About the Team The Monetization team is a new cross-functional group working across engineering, product, research, and design to build the foundational systems that will help OpenAI scale access to intelligence responsibly. Our mission is to develop user-first, privacy-preserving monetization products—including next-generation ads experiences—that strengthen user trust, unlock economic opportunity, and support OpenAI’s long-term innovation. Monetization plays a critical role in enabling OpenAI to continue pushing the boundaries of AI capabilities while ensuring the benefits of AGI are broadly shared. We believe monetization must be aligned with user value, uphold rigorous privacy and safety standards, and sustain a healthy ecosystem of developers and businesses. This team operates in a greenfield environment and moves quickly through prototyping, experimentation, and iterative deployment. We partner closely with Product, Design, and Research to bring research breakthroughs into real-world systems at global scale. About the Role We’re looking for an experienced Software Engineer to help build the machine learning infrastructure that powers OpenAI’s monetization and ads systems. In this foundational role, you’ll design and develop the platform layer that enables teams to build, train, deploy, serve, monitor, and continuously improve machine learning models used across advertising and monetization products. You’ll work across the full ML lifecycle, from large-scale data pipelines and feature infrastructure to training systems, model serving, experimentation platforms, and monitoring frameworks. The systems you build will support high-throughput, low-latency advertising workloads while maintaining strict standards for reliability, privacy, security, and performance. This role sits at the intersection of machine learning systems, distributed infrastructure, and monetization, offering the opportunity to shape the core platforms that help translate model innovation into m

awsrestmachine learning
View job →

About the role We’re looking for an engineering manager to lead a team building software systems that detect and prevent harmful misuse of frontier AI models—before incidents occur. This is a builder’s role: you’ll lead engineers shipping production services, detection pipelines, and mitigation mechanisms that protect frontier model integrity and reduce high-severity misuse risk. While this work intersects with frontier model development, security and risk, we’re explicitly seeking someone with a software engineering foundation who is comfortable building reliable systems that can operate at billions of users scale. In this role you will: Lead a team of software engineers building detection + mitigation systems for frontier model misuse, with an emphasis on model IP protection / distillation detection and emerging risk surfaces from autonomous agents. Set the technical roadmap and execution strategy: prioritize, design, ship, iterate, measure impact. Build production systems: services, pipelines, tooling, instrumentation, and automation that scale with frontier model usage. Partner deeply with Research and Product to translate evolving model capabilities into concrete tests, signals, and mitigations that can be deployed at scale. Drive strong engineering fundamentals: architecture, reliability, monitoring, performance, and operational excellence. Hire and grow an exceptional team across backend, data systems, and applied ML engineering domains as needed. Anticipate what breaks at scale as agentic workflows become more capable. You might thrive in this role if you: Experience building systems in adversarial, fast-evolving environments Are comfortable with ambiguity and novelty Have experience adjacent to security (e.g., abuse prevention, fraud, integrity, platform defense, auth/identity, malware/spam, adversarial environments) Communicate clearly and build trust quickly with senior stakeholders—pragmatic, collaborative, and calm under scrutiny. Significant experience

awsrestai
View job →
S
Stripe
📍 Toronto• Full-time
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Our Applied ML team aims to reform how our users interact with Stripe. We are doing so by (a) automating the easy tasks, and (b) assisting our users in the difficult tasks. Some examples include helping our users resolve issues with Stripe faster or making it easier for our users to sign up and navigate Stripe. We are using the latest LLMs as well as fine-tuning our own models. We're an end-to-end team going from ideas to models to shipping in production. You can learn more about our team’s work from this recent talk . What you’ll do As a machine learning engineer, you will be responsible for analyzing opportunities, proposing ideas, training & evaluating ML models, running experiments, and deploying everything to production. You will also have the opportunity to contribute to and influence ML architecture at Stripe as well as be a part of a larger ML community. Responsibilities Our team operates fluidly and here are some problems you may tackle: How do we evaluate a system offline & online? How do we improve performance to match (and beat) humans? How do we ensure model quality doesn’t degrade online? Does fine-tuning an LLM give us better performance? What are the right OSS and in-house platforms we should invest in? And in the process you will: Develop pipelines and automated processes to train and evaluate models in offline and online environments Integrate ML models into production systems and ensure their scalability and reliab

machine learningaigo
View job →
B
Brex
📍 Sao Paulo• Full-time
16 days ago

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Data at Brex The Data organization develops infrastructure, statistical models, and products using financial data. Our Scientists and Engineers work together to make data —and insights derived from data — a core asset across the company. Our work is ingrained in Brex’s decision-making process, in the efficiency of our operations, in our risk management policies, and in the second-to-none experience we provide our consumers. What You’ll Do Our Data Scientists are responsible for the entire model development lifecycle, from conception with stakeholders, through model development and productionization, to following through to see that the desired business impact is achieved — including circling back with stakeholders to make product or strategic decisions. Responsibilities Drive Data & AI solutions from inception to deployment to efficiently manage risk and/or improve customer experience. Be responsible for the full machine learning

pythonsqlmachine learning
View job →
G
Graphcore
📍 Austin• Full-time
16 days ago

1418 Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. The Electrical Engineer will play a pivotal role in designing innovative hardware systems for AI/ML applications. We are seeking a motivated Electrical Engineer with 2–4 years of experience in schematic capture, PCB design, and server hardware development. This role includes close collaboration with multiple partners to drive designs from concept through mass production. The ideal candidate is comfortable working across organizational boundaries, ensuring design quality, manufacturability, and on-time delivery in a fast-paced environment. Responsibilities: Develop and maintain electrical schematics for various printed circuit assemblies Design and layout multilayer PCBs Work closely with partners to review designs, provide technical guidance, and ensure alignment with system requirements Drive design for manufacturability (DFM), design for assembly (DFA), and design for testability (DFT) Support server subsystem integration Participate in design reviews Support prototype builds, board bring-up, debugging, and validation Track and resolve design issues, including root cause analysis and corrective actions Ensure proper documentation, revision control, and engineering change management (ECO/ECN processes) Requirements: Bachelor’s degree in electrical engineering 2–4 years of experience in schematic capture and PCB design

gitaisem
View job →
C
Cohere
📍 Toronto• Full-time• From £215K/yr
1mo ago

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the role. We’re building the next generation of agentic AI infrastructure at Cohere. This team sits at the intersection of ML systems, distributed infrastructure, and developer experience, creating the platform that powers autonomous AI agents at scale. You’ll work on hard, forward-looking problems with few established patterns, including secure code execution, agent state management, model routing, identity and authentication, and resource management for long-running agent workflows. This role is a strong fit for someone who combines systems depth with ML intuition. You should be comfortable building reliable infrastructure, thinking through distributed systems tradeoffs, and understanding how emerging agentic capabilities shape platform design. What you’ll work on. Secure execution environments for agent-generated code Identity, authentication, and trust boundaries for agents Model routing and orchestration across different model types and environments Rate limiting, quotas, and resource management for agent workflows State management, memory, and filesystem abstractions for agents. In this role you will: Turn emerging M

kubernetesgitrest
View job →
P
Pinterest
📍 Palo Alto• Full-time• Remote• From $114.3K/yr
1mo ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . This role focuses on advancing the science and systems behind ML measurement, feature understanding, and causal inference at scale. The work spans areas such as production feature importance platforms, observational causal estimation in Pytorch, large-scale proxy metric development, and data-driven approaches to ML infrastructure efficiency. We're looking for an enthusiastic individual contributor to perform high-impact technical work across this space. This person will drive foundational innovations, own the end-to-end design of production ML systems, establish rigorous methodological standards, and partner cross-functionally to turn successful research into durable platform capabilities that raise the ceiling for the entire ML organization. What you’ll do: We are looking for an experienced and highly capable Data & Applied Scientist

REMOTEpythonawsrest
View job →
T
16 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent’s AI Models team and work at the layer most ML engineers never see: bringing advanced models to life on custom AI hardware. You’ll own real workloads end‑to‑end including porting, tuning, and validating LLMs and vision models on our accelerator, and chasing down every last millisecond and percentage point of accuracy. This role is for people who love the craft of ML engineering and want their work to matter at silicon scale, not just behind another API. This role is hybrid , based in Cyprus. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Bring up, run, and debug modern ML models (e.g., transformers) using PyTorch or TensorFlow. Analyze model behavior and performance, and identify bottlenecks across the stack. Improve efficiency, correctness, and scalability of model execution in real systems. Work closely with compiler, kernel, and hardware teams to drive performance and system-level improvements. Help translate state-of-the-art model architectures into production-grade, high-performance deployments. What We Need Strong experience building and working with ML models in PyTorch or TensorFlow. Strong understanding of mod

awsaic++
View job →
R
1mo ago

Join us in building the future of finance. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. ABOUT THE TEAM + ROLE We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The AI R&D team is at the core of Robinhood's product intelligence. Our mission is to build and scale high-impact models that power personalization, search, social feeds, fraud detection, and risk management for millions of Robinhood users. We operate as a cross-functional partner to growth, product, and data engineering—translating complex financial data into intelligent systems that make Robinhood smarter for every customer. We move fast, raise the bar, and care deeply about building things that matter. If you've ever wanted to solve personalization problems no one else has cracked—in one of the most data-rich, regulated industries on the planet—this is the team for you! As a Staff Machine Learning Engineer on the AI R&D team, you will own the design and delivery of sophisticated personalization and recommendation systems that directly shape what millions of users see and do on the Robinhood platform. You'll be a technical anchor on a growing, high-caliber team — collaborating with product, data engineering, and fellow ML engineers to take ambitious ideas from zero to one and into production at scale. You'll help define the team's technical direction, mentor engine

pythonvueaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in cloud-native environments, in close partnership with Amazon and internal teams across Codex, Research, Safety Systems, and Applied. The team is focused on bringing core developer and enterprise capabilities into cloud-native environments, including cloud-hosted Codex, model customization / post-training as a service, and new stateful runtime environments for agentic workloads. This work sits at the intersection of production ML systems, developer platforms, model behavior, and large-scale infrastructure. About the Role We’re looking for a backend engineer who can quickly understand OpenAI’s models, products, and systems, then adapt first-party deployments for other cloud platforms. You’ll build backend services, APIs, SDK integrations, authentication flows, and cloud service infrastructure that let developers use OpenAI capabilities in the cloud environments where they already build. This role involves working across teams, sometimes embedded with partner product groups, to ship products quickly and across multiple platforms at the same time. It’s a strong fit for engineers who have built developer tools, especially AI-powered tools, communicate clearly across technical boundaries, and can shape architectures that support different deployment models; experience building cloud services is a strong plus. In this role, you will: Build backend and infrastructure systems that extend OpenAI’s API platform into cloud-native environments, like AWS. Design and ship cloud-contained products that allow customers to use OpenAI capabilities while keeping workloads and data within cloud environments. Help stand up cloud-hosted Codex experiences powered by the OpenAI Responses API. Build the infrastructure and runtime abstractions

typescriptpythonaws
View job →

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu

pythonkubernetesci/cd
View job →

The Anyscale Technical Program Management (TPM) team is expected to play a critical role executing high impact programs while continuously improving processes to sustainably grow and increase the effectiveness of the Tech organization spanning Design, Engineering and Product teams. As a Technical Program Manager focused on Anyscale’s core product solution, you’ll help lead complex application development in service of enhancing our product platform. In this role, you'll support and help scale the technical solutions that make Anyscale’s products and services possible. As part of the overall development life cycle you’ll plan requirements, identify risks, manage schedules, and communicate clearly with project stakeholders on complex projects with significant bottom line impact. Your Program management contributions will span prioritization, planning of projects and features, stakeholder management, tracking of external commitments while contributing to the organization's technical culture by highlighting and espousing best practices. You’ll learn and grow alongside talented teammates who share your commitment to excellence and appetite for innovative problem-solving. This is a rare opportunity to join in the leadership of a team that will be responsible for building a successful commercial ML/AI oriented solution from the ground up! As part of this role, you will: Help us build, track and ship our Commercial / OSS product Will work closely with the software development and product teams to deliver high quality, scalable products used by customers around the world Collaborate with the product teams and align all the stakeholders to assemble project teams, assign responsibilities, identify appropriate resources needed, and develop schedules to ensure timely completion of projects by meeting project milestones Assess risks, anticipate bottlenecks, provide escalation management, make tradeoffs, balance the business needs versus technical constraints and encourage risk ta

machine learningaigo
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe processes over $1T in payments volume per year, which is roughly 1% of the world’s GDP. The tremendous amount of data makes Stripe one of the best places to do machine learning. The ML Infra team builds services and tools that power every step in the ML lifecycle, including data exploration, feature generation, experimentation, training, deploying, serving ML models, and building LLM applications. With the phenomenal developments happening in the field of AI, we are positioned to accelerate the adoption of AI/ML across all parts of the company by building highly scalable and reliable foundational infrastructure. What you’ll do You will work closely with machine learning engineers, data scientists, and product engineering teams to enable seamless end-to-end experience in building solutions across data, analytics, and AI/ML platforms. You will build the next generation of ML Infra services and major new capabilities that substantially improve ML development velocity and MLOps maturity across the company. Responsibilities Designing and building scalable, reliable, and secure services for notebooks, ML model training, experimentation, serving, and LLM applications across multiple regions. Creating services and libraries that enable ML engineers at Stripe to seamlessly transition from experimentation to production across Stripe’s systems. Working directly with product teams and ML engineers to improve their day-to-day pr

restmachine learningai
View job →

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Role Overview & Key Responsibilities This is a high-leverage leadership role that spans architecture, execution, and org-building, and will shape the direction of our AI / ML initiatives at Ema. We are seeking an AI / ML technical leader who can take a vision and build it. As a Principal ML Engineer at Ema, you will be a senior technical leader responsible for shaping the machine learning roadmap, architecting large-scale ML systems, driving innovation, and ensuring our mixture of expert models (LLM + SLM + Custom Model) is accurate and performant at scale. You will collaborate across teams (research, product, infra, data, etc.), mentor senior engineers, and influence strategy and execution at company-wide levels. Responsibilities Lead the technical direction of GenAI and agentic ML systems that power enterprise-grade AI agents — spanning reasoning, retrieval, tool use, and integrations across various SaaS products. Architect, design, and implement scalable production pipelines for model training, fine-tuning, retrieval (RAG), agent orchestration, and evaluation — ensuring robustness, latency efficiency, and continuous learning. Define and own the multi-year ML roadmap for GenA

pythonjavamachine learning
View job →
🔔

Get new ml platform engineer jobs by email

Daily job updates · Unsubscribe anytime