At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build the future of data. Join the Snowflake team. The Snowflake Machine Learning Platform team’s mission is to enable customers to bring their machine learning and deep learning workloads to Snowflake. Our customers want to build powerful models with the ever-increasing data in Snowflake but face several challenges including infrastructure optimizations, orchestration, performance, and security. The team aims to solve these challenges by building highly integrated platform solutions that are simple, secure, and enable end-to-end ML workflows. We are on an early journey to build the most scalable machine learning and data platform without sacrificing the benefits of a single platform and governance. We are looking for outstanding technical leaders who will join our ML Platform team to build the next-generation platform and play a pivotal role in this journey by understanding Snowflake’s core platform architecture and evolving it to enable state-of-the-art machine learning and LLM workloads. Join us to define strategies, set technical directions, design and execute, engage and deliver innovation, and unlock the power of AI for thousands of enterprise customers. This position is based in Menlo Park, CA, and Bellevue, WA. RESPONSIBILITIES : Help define and own the roadmap, wor
Jobiba hiring network
Ml Platform Engineer Jobs
832 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current ml platform engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About Okta Okta is an enterprise grade identity management service, built from the ground up in the cloud and delivered with an unwavering focus on customer success. With Okta you can manage access across any application, person, or device. Whether the people are employees, partners, or customers, or the applications are in the cloud, on premises, or on a mobile device, Okta helps you become more secure, make people more productive, and maintain compliance. The Okta service provides directory services, single sign-on, strong authentication, provisioning, workflow, and built in reporting. It runs in the cloud on a secure, reliable, extensively audited platform and integrates deeply with on premises applications, directories, and identity management systems. About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Softwa
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response. You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure. RESPONSIBILITIES Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s. Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work. Foster a culture of ownership, accountability, and continuous improvement. Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate th
About Scale Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust. About the Team Applied Intelligence Systems (AIS) is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. AIS spans multiple workstreams — agent evaluation and oversight, orchestration and tool-use infrastructure, model and systems optimization, and applied research on new agent capabilities — and this role is not scoped to any single one of them. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like. About the Role As a Staff Machine Learning Research Engineer, you will operate across the full breadth of AIS’s technical needs — wherever the hardest ML problem in agentic AI happens to be that quarter. This could mean training and fine-tuning models, designing evaluation and observability systems, building improvement loops from production data, prototyping novel agent architectures, or designing internal systems and tooling that boost productivity across teams. You’re not tied to one team’s roadmap; you’re expected to move to where the technical leverage is highest, and t
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative AI platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale our product in a fast-paced environment. The ideal candidate will have a strong understanding of software engineering principles and practices, as well as experience with large-scale distributed systems. You will be responsible for owning large new areas within our product, working across backend, frontend, and interacting with LLMs and ML models. You will solve hard engineering problems in scalability and reliability. You will: Own large new areas within our product Work across backend, frontend, and interacting with LLMs and ML models Deliver experiments at a high velocity and level of quality to engage our customers Work across the entire product lifecycle from conceptualization through production Be able, and willing, to multi-task and learn new technologies quickly Ideally you'd have: 7+ years of full-time engineering experience, post-graduation Experience scaling products at hyper growth startups Experience tinkering with or productizing LLMs, vector databases, and the other latest AI technologies Proficient in Python or Javascript/Typescript, and SQL Experience with Kubernetes Experience with major cloud providers (AWS, Azure, GCP) Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval
SPECIFIC JOB RESPONSIBILITIES Pipeline Management: Maintain high-throughput streaming pipelines to ingest logs from various sources (Firewalls, Cloud, Endpoints) to a central destination. Log Normalization: Write parsers to convert raw, messy logs into standard schemas (e.g., OCSF or ECS) for consistent querying. Cost Optimization: Implement routing logic to send "high-value" data to the SIEM and "bulk" data to low-cost Object Storage (Data Lake). Data Preparation: Clean and structure data to enable AI/ML detection models and advanced analytics. EXPERIENCE REQUIRED Data Engineering: Proficiency in Python (for ETL) and SQL (for complex querying). Streaming Tech: Experience with Message Queues (e.g., Kafka, Pub/Sub) and stream processing concepts. Log Handling: Mastery of Regex and log parsing strategies for standard formats (Syslog, CEF, JSON). Storage Architecture: Understanding of Data Lake principles (Parquet/Avro formats) vs. Data Warehouses. QUALIFICATIONS, SKILLS, & KNOWLEDGE Experience with Vector Databases for storing embeddings. Knowledge of Log Observability/Routing tools (middleware that routes logs). Familiarity with Big Data frameworks (e.g., Spark, Flink). PROFESSIONAL DEVELOPMENT EXPECTATIONS Ability to embrace Clearwater's CLEAR core values (Commitment to Client Success, Lead with Accountability, Integrity & Collaboration, Excellence in All That We Do, Advance Colleague Success, Respect & Transparency) and culture. The base salary range for this role is 35,000- 45,000]. Base salary is part of our total rewards package which also includes the opportunity for merit-based salary increases, eligibility for our 401(k) plan, medical, dental, vision, life and disability insurances and leaves provided in line with your work state. Our robust time-off policy includes flexible paid time off, 11 paid holidays, and paid sick time. Total compensation, including base salary to be offered, will depend on elem
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The Role We are looking for an Engineering Leader to manage and scale multiple product lines in the Voice, BPO, and Workforce Management space. This is a high-impact leadership role that sits at the intersection of real-time voice systems, operations research, and data-intensive platform engineering. You will report directly to the Head of Engineering and own the engineering organization that builds the infrastructure powering Ema’s Voice AI Employees, Agent QA, auto-learning pipelines, rich analytics, and workforce optimization capabilities — all operating as scalable, multi-tenant systems deployed across global geographies. You will collaborate with Product, ML/AI, and Go-to-Market teams to translate customer needs into production systems that handle high volumes of voice data, deliver real-time insights, and continuously improve through automated learning loops. As the owner of multiple product lines, you will balance roadmap priorities across Voice, BPO operations, and WFM (work force management) — ensuring each product evolves cohesively while meeting distinct customer needs. What You Will Do Scalable Multi-Tenant Systems Architect and build multi-tenant systems that serve ent
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacart’s AI Productivity team builds AI-powered platforms and tools that help our engineers and operators move faster, reduce toil, and deliver higher-quality experiences for customers, shoppers, retailers, and brand partners. As a Senior Software Engineer on this team, you will design, build, and operate production systems that bring large language models and intelligent automation into everyday workflows across Instacart. You’ll partner closely with product, developer platform, ML platform, security, and data teams to ship reliable, secure, and measurable solutions that improve developer velocity and operational efficiency. You’ll join a collaborative group of approximately 10 engineers who value ownership, iteration, and pragmatic problem solving. If you enjoy rolling up your sleeves, navigating ambiguity, and turning cutting-edge AI into real, scala
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. We are looking for a Senior Software Engineer specializing in Machine Learning to join our Revenue ML team at Discord. This team partners with our revenue product groups, focusing on both consumer revenue and our emerging Ads initiative. This role will specifically contribute to our Ads ML efforts, helping to build and scale ML capabilities in areas such as ads measurement, targeting, and delivery ranking. As part of this team, you will play a critical role in developing foundational ML models that enhance ad relevance, optimize performance, and drive revenue. This is a unique opportunity to work on an early-stage Ads ML platform and have a direct impact on the business's success. Our tech stack includes Python, ML frameworks like PyTorch and TensorFlow, large-scale data infrastructure, and real-time ad-serving technologies. What You'll Be Doing: Design, develop, and deploy machine learning models for ads targeting and ranking. Develop sophisticated ML solutions such as identity graph to enhance ad targeting. Build and optimize ad ranking models to serve the most effective ads based on campaign objectives (e.g., app installs, link click). Improve ads targeting and ranking by leveraging both on-platform and off-platform signals. Collaborate cross-functionally with product, engineering, and business teams to define and execute on the Ads ML roadmap. Scale our ML infrastructure to support an increasing number of concurrent ad campaigns while ensuring low-latency decision-making. Drive research and implementation of state-of-the-art ML techniques in the field of online advertising. What You Should Have: 5+
Leidos is looking for our next TS/SCI-cleared Elastic Search Engineer to join a high-energy team building and deploying a cutting-edge technology stack to support our client’s mission to centralize and standardize Tasking, Collection, Processing, Exploitation and Dissemination (TCPED) of Open Source Intelligence (OSINT) across the DoD and IC enterprise. We integrate off the shelf and newly development software to sustain and enhance the TCPED platform. We leverage cloud-based computing, artificial intelligence (Al), machine learning (ML) and cross-domain transfer systems to provide cutting edge data exploitation, enrichment, triage, and analytics capabilities to Defense and Intelligence Community members. DTP advances the state of the art in mission-focused big data analytics tools and micro-service development spanning the breadth of Agile sprints to multiyear research and development cycles. As an Elastic Search Engineer, you’ll be a member of our platform engineering team and help develop, deploy and maintain nosql databases as foundational elements of our microservice eco-system using a Kubernetes as foundational platform. You’ll also support the adoption of GitOps best practices across cross-functional engineering teams. In this fast-paced environment, you’ll collaborate closely with systems engineering, architecture, development, security, operations, and integrations teams. Work is conducted on-site at our client location in Bethesda, MD. Key Responsibilities Include: Deploy, triage, debug, and maintain production class databases like Elasticsearch and Redis Design and support database configuration management strategies across air-gapped network fabrics Partner with Systems Engineers to architect solutions for new capabilities Contribute to operational monitoring capabilities to provide proactive system notifications Contribute technical input to engineering documentati
Leidos is looking for our next TS/SCI-cleared Elastic Search Engineer to join a high-energy team building and deploying a cutting-edge technology stack to support our client’s mission to centralize and standardize Tasking, Collection, Processing, Exploitation and Dissemination (TCPED) of Open Source Intelligence (OSINT) across the DoD and IC enterprise. We integrate off the shelf and newly development software to sustain and enhance the TCPED platform. We leverage cloud-based computing, artificial intelligence (Al), machine learning (ML) and cross-domain transfer systems to provide cutting edge data exploitation, enrichment, triage, and analytics capabilities to Defense and Intelligence Community members. DTP advances the state of the art in mission-focused big data analytics tools and micro-service development spanning the breadth of Agile sprints to multiyear research and development cycles. As an Elastic Search Engineer, you’ll be a member of our platform engineering team and help develop, deploy and maintain nosql databases as foundational elements of our microservice eco-system using a Kubernetes as foundational platform. You’ll also support the adoption of GitOps best practices across cross-functional engineering teams. In this fast-paced environment, you’ll collaborate closely with systems engineering, architecture, development, security, operations, and integrations teams. Work is conducted on-site at our client location in Bethesda, MD. Key Responsibilities Include: Deploy, triage, debug, and maintain production class databases like Elasticsearch and Redis Design and support database configuration management strategies across air-gapped network fabrics Partner with Systems Engineers to architect solutions for new capabilities Contribute to operational monitoring capabilities to provide proactive system notifications Contribute technical input to engineering document
Staff Software Engineer - Testing & Automation Exceptional software engineering is challenging. Amplifying it to ensure that multiple teams can concurrently create and manage a vast, intricate product escalates the complexity. As a Staff Engineer within the Verification Platform team at Sumo Logic, you will drive the implementation and optimization for our verification platform as well as the modernization of our CI/CD pipelines. Your mission is to develop and sustain automated tooling for all testing, verification, and functional requirements, leveraging AI reasoning and machine learning models to predict and prevent delivery issues, while integrating advanced security validation and non-functional requirements into our delivery lifecycle. You will contribute significantly to establishing automated delivery pipelines, empowering autonomous teams to create independently deployable services, and progressing Sumo Logic’s internal Platform-as-a-Service. This role sits at the intersection of Platform Engineering, Quality Engineering, DevSecOps, and Developer Productivity, helping teams deliver secure, reliable, and independently deployable services at scale. Responsibilities Strategy & Leadership: Drive technical direction and design for a modern Quality Engineering platform, driving the adoption of AI reasoning for enhanced automation of all testing, verification, and functional requirements. Pipeline Modernization: Lead the modernization of CI/CD pipelines to include automated security validation, compliance checks, and other critical non-functional requirements, with a focus on integrating AI/ML for intelligent pipeline optimization and risk prediction. Framework Ownership: Own the delivery pipeline and release automation framework for all Sumo services, ensuring improvements in developer productivity, deployment frequency, and release reliability. Cross-Team Collaboration: Educate and collaborate with teams during design and development phases to ensur
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With over half a billion rides and counting, Lyft is solving hard problems in a flourishing domain with a lot of data and creative solutions in Marketplace, Mapping, Fraud, Growth and beyond. We're actively building the next-generation Machine Learning (ML) platform for low-cost, ultra-immersive transportation to improve people’s lives using modern ML with peta-byte scale data. Our Machine Learning Engineers are excited to work on these challenging problems and redefine solutions to directly impact various aspects of Lyft's primary business. If you are a student with experience in machine learning workflows, passionate about solving challenging problems using data and working in a dynamic, creative, and collaborative environment, this opportunity is for you! Responsibilities: Contribute to the design, build, train and test of Machine Learning models Write production-level code to convert ML models into working pipelines Partner with Product Managers, Data Scientists, and fellow ML Engineers to frame Machine Learning problems within the business context Analyze experimental and observational data, communicate findings to support decisions Participate in code and spec reviews to ensure code quality and distribute knowledge Experience: Currently pursuing a Bachelor's, Master's, or PhD degree in Computer Science or a related technical field from a university in Canada (required) , with a graduation date between December 2027 and Summer 2028 (required). For any candidates who are master's students who worked between their bachelor's and master's programs: candidates should also have less than 2 years of relevant full-time work experience Available during Summer 2027 for the internship in Toronto Good understanding and knowledge of ML libraries like scikit-learn, Tensorflow, PyTorch, Keras, MXNet, et
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. With a billion rides per year and counting, Lyft is solving hard problems in a rapidly growing domain with a lot of data and creative solutions in Rider, Driver, Marketplace, and beyond. While traditional approaches to optimization and problem decomposition are sufficient to disrupt transportation, building a next-generation platform for low-cost, ultra-immersive transportation to improve people's lives warrants modern ML utilizing petabyte-scale data. Our highly motivated Machine Learning Engineers work on these challenging problems and define solutions to directly impact various aspects of our core business. The Fulfillment group, within the Marketplace at Lyft, is responsible for determining what inventory can be reliably offered for a given rider session and fulfilling rider requests. The group comprises several sub-teams that generate feasible offers for riders, match rider requests with drivers, and maintain a distributed state machine to track rides and drivers from request through completion. We are seeking a Machine Learning Engineer to join the Fulfillment team and lead the design, development, and deployment of state-of-the-art machine learning systems. This role requires a strategic thinker who can balance high-level system architecture with hands-on technical implementation. You will collaborate across teams to shape the future of ride-sharing by leveraging machine learning and data science. Responsibilities: Design, build, and deploy machine learning models for real-time applications, including translating state-of-the-art research into production-ready solutions Design and implement feature pipelines, model training workflows, and serving infrastructure using Lyft's ML platform Evaluate ML system performance against business KPIs, run experiments, and drive continuous model improvement
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is looking for hands-on, passionate people who want to join a high energy and growing team to make a difference in customers’ lives and who want to be on the forefront of digital innovation that aims to reinvent what a pharmacy and a health care company can be in the digital world. Currently, we are seeking a Staff Software Engineer – Search / AI who as a Senior technical leader, be responsible for driving architecture, design, and delivery of scalable, cloud-native platforms built on microservices architecture and AI capabilities. This role combines deep hands-on engineering with strategic leadership to build intelligent, distributed systems. The right candidate will be a strong analytical thinker and be able to simplify complex problems, processes or projects into component parts explore and evaluate them systematically. We love to collaborate and help each other and we want someone to share that ideology. Expectations for the Role Drive enterprise architecture and technical strategy with strong focus on microservices-based design and AI platform engineering Design and develop highly scalable microservices architectures, including APIs, domain-driven services, and event-driven systems Lead the development and integration of AI/ML solutions, including LLMs, Retrieval-Augmented Generation (RAG), and agentic frameworks Develop sc
Get new ml platform engineer jobs by email
Daily job updates · Unsubscribe anytime