Jobiba hiring network

Ml Platform Engineer Jobs

832 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current ml platform engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs. If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact. What You’ll Work On Build and own the training framework responsible for large-scale LLM training. Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing). Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100). Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics. Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training. Investigate and res

dockerkubernetesgit
View job →
F
Fin
📍 Ireland• Full-time
1mo ago

Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's AI Group is responsible for defining new ML products, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. We are extremely product-focussed. Our team of 50+ ML scientists, ML engineers, designers and researchers works in partnership with other teams across the whole company. We move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, novel applications of transformer neural networks and fine-tuned LLMs. We test and measure the real customer impact of everything we deploy. We plan to double in size, and are on the lookout for an experienced manager to manage a team of highly-performing ML scientists. Engineering managers in Fin are not just people

sqlmachine learningai
View job →
D
Datadog
📍 Massachusetts• Full-time• From $234K/yr
1mo ago

The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging Large Language Models (LLMs) and generative AI. We provide robust, scalable observability for AI workloads, including drift detection and model evaluation, and behavior tracing, enabling customers to ship AI with confidence. As a Staff Engineer, you’ll lead the development of new features and foundational capabilities within Datadog’s LLM Observability product. You will shape product direction, drive experimentation, and apply your deep understanding of both AI systems and software engineering to solve open-ended problems in the fast-moving AI landscape. Your work will directly impact how our customers monitor, troubleshoot, and optimize LLM-based applications in production. Join us in building the foundational tools that make AI systems observable, understandable, and reliable in the real world. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Drive design and implementation of LLM observability features. Ideate, prototype, and scale new product features to provide insights and drive improvements for generative AI systems Work cross-functionally with other eng teams, product, UX, and applied science to iterate fast and find product-market fit Develop and extend tools for tracing, evaluating, and debugging LLMs Influence architecture decisions and mentor engineers to build resilient, high-performance systems Stay close to customer pain points and use those insights to guide product and engineering priorities Stay current with industry trends and advancements in machine learning and observability, driving innovation within the team Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or r

machine learningaigo
View job →
BA
Bolna AI
📍 India• Full-time
1mo ago

At Bolna, we’re building tools that change the way teams leverage Voice AI. We’re looking for a Founding Machine Learning Engineer to own the end-to-end lifecycle of building, evaluating, deploying, and improving models that power millions of production conversations. This is a high-impact, high ownership role where you won’t just work on Bolna’s ML stack—you’ll help build the foundation it scales on. Our team includes IIT alumni with experience at Bain, Atlassian, Uber, Zomato, and LinkedIn, and is backed by leading investors. Responsibilities: Build the data engine - Design pipelines to source and clean conversational voice data across Indian languages, accents, and telephony conditions. Fine-tune models that ship - Fine tune and train models to improve accuracy, speed, and reliability across different use-cases. Define what "good" means - Build evaluation datasets and benchmarks for transcription accuracy, voice naturalness, interruption handling, latency, and end-to-end conversation quality. Set up human-in-the-loop pipelines to capture subjective quality at scale. Ship to production - Work with the engineering team to deploy models into a latency-sensitive, high-volume system. Monitor performance in the wild, debug regressions, and iterate fast. Required Skills: 3+ years of hands-on ML experience with deep practical real-world experience in training models. Strong Python and PyTorch fundamentals with exposure in distributed training, and modern fine-tuning techniques (LoRA, QLoRA, DPO, RLHF, etc.). Training data as a first-class problem. Experience designing data pipelines from collection, cleaning, labeling, deduplication, augmentation and treating data quality as a core engineering discipline. Rigorous about evaluation. You know that "looks good in a demo" is not a benchmark. You build the evals before you trust the model. Speech model experience is a plus with real-time / streaming inference experience where you would have contributed to latency optimization

pythonmachine learningai
View job →
S
16 days ago

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Role Overview As our lead AI/ML Engineer , you will design, build, and scale the intelligence layer powering our next-generation security products. You will turn complex datasets—including configurations, security alerts, and raw network logs—into production-ready AI capabilities that drive automated analysis, reasoning, and intelligent recommendations. Key Responsibilities Architect AI Systems: Design and deploy robust Retrieval-Augmented Generation (RAG) pipelines, agentic workflows, and LLM applications tailored to parse and reason over security telemetry and unstructured logs. Model Development & Tuning : Train, fine-tune, and evaluate ML/DL models for pattern matching, anomaly detection, and classification across massive, heterogeneous security datasets. Ensure AI Trust & Alignment : Implement strict guardrails, evaluation frameworks, and safety alignment techniques to ensure all model outputs and recommendations are deterministic, safe, and highly accurate. Establish MLOps: Build and maintain scalable ML pipelines (from data preprocessing to model monitoring in production), collaborati

pythonawsazure
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Our Solution Engineering organization is seeking an AI Specialist who can provide hands-on expertise and support while working with technical decision makers and data scientists to design and architect AI solutions built on the Snowflake AI Data Cloud. This is a strategic role that works closely with cross-functional teams, including product, engineering, and the broader field organization to ensure successful execution and customer adoption of Snowflake’s AI & ML solutions. IN THIS ROLE YOU WILL GET TO: Be the technical expert in the room that positions Snowflake’s AI and ML features and value to technical stakeholders at Snowflake’s customers across the Americas. Partner with Snowflake account team teams and customer champions to scope and drive POCs to success and technical wins that prove the value of Snowflake’s capabilities, including executive readouts and business value cases. Collaborate with Snowflake’s product and engineering teams to influence Snowflake’s AI and ML roadmaps based on customer feedback. Publish content that helps the team and company scale beyond your individual efforts, like blog posts, presentations at conferences, or technical collateral like notebooks and demos. Influence, tailor and maintain Sales Engineering AI and ML selling assets, inc

pythonawsazure
View job →

Step into a mission where what you do truly matters! At Leidos, innovation is at the heart of everything we do. Powered by a team as diverse as it is talented, we're driven by a shared passion for delivering bold solutions that fuel our customers' success. We believe in empowering our people, giving back to our communities, and leading with sustainability. Every action we take is grounded in integrity and a steadfast commitment to doing what’s right—for our customers, our teams, and the world around us. Our Mission, Vision, and Values aren't just words—they're the compass guiding our journey toward a brighter future. If this sounds like the kind of environment where you can thrive, keep reading! Leidos is thrilled to share an exciting opportunity for a mission-focused Intelligence Analyst (AI/ML) at Joint Base Langley–Eustis in Hampton, VA! This is a chance to directly shape the next generation of Air Force intelligence professionals through cutting-edge AI/ML, data analytics, and automation training. Candidates must currently hold an active TS/SCI clearance . What You’ll Do: Help empower the ACC A2 mission by delivering high-impact instructional support: Assist in providing instructional support services that will allow the A2 staff to achieve its mission: Independently deliver and maintain formal classroom and virtual instruction on tool-agnostic AI/ML concepts, Large Language Models (LLMs), Robotic Process Automation (RPA), and advanced data analytics, ensuring established courseware remains current with standard DoD and IC applications. Execute tiered role-based training, adapting delivery to meet varying student proficiency levels—ranging from foundational AI/ML user adoption and awareness to intermediate prompt engineering and automated workflow creation. Instruct personnel on the underlying methodologies of AI/

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are seeking experienced professionals with a strong background in Artificial Intelligence, Machine Learning, and Cloud Architecture to join our Services Delivery team to help create exciting new offerings and capabilities for our customers! In this strategic role, you will help customers expand their use of the Snowflake Data Cloud to bring AI/ML pipelines from ideation to full production. Leveraging Snowflake’s native features and extensive partner ecosystem, you will advise clients on best practices for scaling production-ready workloads. You will design tailored AI/ML solutions, coordinate closely with customer teams and Systems Integrators, and provide the technical leadership and oversight needed to ensure successful outcomes. AS A PRINCIPAL SOLUTIONS ARCHITECT AT SNOWFLAKE, YOU WILL: Be a technical expert on all aspects of Snowflake in relation to the AI/ML workload and provide customers with best practices given Snowflakes technology stack. Work with customers to understand their AI/ML use case, discover key requirements, and architect a Snowflake-centric solution to be delivered by Services Delivery. Understand how to build, deploy and AI and ML pipelines using Snowflake features and/or Snowflake ecosystem based on customer requirements. Work hands-on where neede

pythonjavasql
View job →
🔔

Get new ml platform engineer jobs by email

Daily job updates · Unsubscribe anytime