Jobiba hiring network

Senior Machine Learning Operations Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior machine learning operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

SA
Scale AI
📍 San Francisco• Full-time• From $302.4K/yr
15 days ago

Director of Engineering, Physical AI Role Overview The Director of Engineering will report to the General Manager of Physical AI, and will be responsible for leading a multi-disciplinary engineering organization. In this senior leadership role, you will own the execution of the Physical AI Data Engine — the platform powering the next generation of Physical AI/Embodied AI. You will collaborate closely with Operations and GTM to guide product direction and help solve the data bottleneck that stands between today's robotics research and real-world deployment. This role requires significant ownership in a fast-paced environment and you will motivate internal teams to set the pace for business growth. Travel will come into play. Key Responsibilities: Set and drive the technical vision across data collection infrastructure, teleoperation systems, ML training pipelines, model evaluation frameworks, annotation tooling, and research Lead a multidisciplinary engineering organization—spanning engineering managers, software engineers, ML engineers, and ML research scientists—while designing the organizational structure, talent strategy, and culture required to scale rapidly without compromising on quality or strategic alignment Maintain exceptional technical and operational excellence by deeply understanding team deliverables, asking incisive questions, identifying slipping standards early, and knowing precisely when to step in Drive cross-functional alignment across Engineering, Operations, and GTM on platform architecture, release processes, and shared priorities Collaborate with researchers and clients to architect and deliver scalable, production-grade data infrastructure tailored for complex robotics workloads Required Qualifications: Bachelor's degree in Engineering, Robotics, Computer Science, or a related technical field 8+ years of engineering experience in fast-paced environments, including 4+ years direct people management demonstrated history of recruiting, mentorin

typescriptpythonaws
View job →
P
Plaid
📍 San Francisco• Full-time
1mo ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are the first line of defense against fraud and abuse on the Plaid platform. Our mission is to ensure the safety and integrity of our platform for consumers and customers. As a Fraud and Abuse Operations Analyst , you will be responsible for responding to fraud and abuse events, investigating claims, and triaging incidents. We also partner with product and engineering teams to inform and improve fraud mitigation strategies. Responsibilities: Safeguard Plaid's Platform: Participate in the abuse on-call rotation, directly protecting our users and customers by responding to and resolving fraud and abuse events. Your timely actions will be instrumental in maintaining trust and security. Drive Investigations and Mitigate Risks: Investigate fraud and abuse claims from diverse sources, partnering with senior teammates on complex cases. Your findings will inform decisions and strategies, directly impacting Plaid's ability to prevent future incidents and minimize financial losses. Proactively perform threat modeling of abuse surfaces and continuously survey external fraud trends, adversary techniques, tooling, and emerging threat vectors Support Incident Response: Help triage and manage fraud and abuse ev

sqlawsmachine learning
View job →
R
Roblox
📍 San Mateo• Full-time• From $280.5K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Every day, Roblox users play over 7M experiences, make over 270M avatar updates, send 6+ billion chat messages and talk for 1+ million hours using voice. Safety is Roblox’s top priority. As part of the Safety Group, the Content and Communications Safety team builds AI features and tools to automate review and moderation of all of these activities on the platform. Areas that the Content Safety Team works on include: Experience reviews The components that make up those experiences (textures, 3D meshes, models, scripts) Avatars, Avatar items Audio, Video Text chat and voice conversations and other content types As a Senior Product Manager on the team, you will work closely with leadership, product, engineering, data science, and operations teams across Roblox. In addition, you will drive critical company-level KPIs that affect how Roblox keeps the platform safe and civil. As we bring new immersive features and content types to Roblox, you will be responsible for leading the product in one of the most challenging and vital roles at Roblox. You will Lead safety features focused on automation of content reviews by leveraging AI Stay one step ahead of malicious actors and their adversarial tactics

awsgitmachine learning
View job →
A
Airbnb
📍 - USA• Full-time• Remote• From $179K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: “ Our real innovation is not allowing people to book a home; it’s designing a framework to allow millions of people to trust one another. Trust is the real energy source that drives Airbnb… ” - Brian Chesky, Airbnb Co-Founder & CEO (2019) Data science is the engine behind Airbnb's most impactful decisions. The Platform Data Science team accelerates product evolution and business outcomes by combining scientific rigor with deep domain expertise, spanning experimentation, machine learning, causal inference and scalable intelligence. We partner closely with product, engineering, policy, and operations teams across Trust to detect and defend against the adversarial behavior that threatens guest and host trust: fraudulent listings and fake inventory, review and content manipulation, account takeover, and other bad-actor activity on the platform. Whether measuring the impact of a new listing integrity defense, modeling risk at the listing or account level, or evaluating the effectiveness of an enforcement policy, our work helps guests and hosts experience an Airbnb that is safer, smarter and more personalized. The Difference You Will Make: This role sits at the heart of some of Airbnb's most consequential data science challenges, where rigorous statistical thinking and applied ML directly shape platform outcomes. You will own high-visibility initiatives that require both technical depth and strong business judgment - work that is visible to leadership and has measurable impact on Airbnb's users and bottom line. A Typical Day: The ideal candidate is a technically exce

REMOTEpythonsqlmachine learning
View job →

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. At CVS Health, Site Reliability Engineering (SRE) is fundamental to delivering the reliable, secure, and scalable technology experiences that support millions of patients, customers, pharmacists, and healthcare professionals every day. Our SRE organization drives operational excellence across critical healthcare and retail platforms through innovation, automation, observability, and engineering best practices. The Executive Director, Site Reliability Engineering serves as the strategic leader responsible for the reliability, resilience, and performance of CVS Health's retail and pharmacy technology ecosystem. This executive will define and execute a comprehensive reliability strategy, oversee large global engineering teams, and establish a long-term vision for observability, automation, and operational excellence across thousands of store locations. Working closely with senior business and technology leaders, the Executive Director will champion modern SRE practices, accelerate incident response capabilities, and deliver real-time operational visibility that enables proactive issue prevention and exceptional customer and patient experiences. Key Responsibilities Strategic Leadership & Vision Define and lead the enterprise-wide Site Reliability Engineering strategy supporting CVS Health's retail and pharmacy operations. Align reliability and operational

awsazuregcp
View job →
G
Gitlab
📍 United States• Full-time• Remote• From $86.4K/yr
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Senior Internal Auditor reporting to the Senior Manager, Technology Internal Audit, you’ll help GitLab assess risk and strengthen controls across a technology landscape that includes multi-cloud infrastructure, artificial intelligence and machine learning systems, and modern development practices. This USA-based role supports our Sarbanes-Oxley Act (SOX) program while partnering with Engineering, IT Operations, Security, and business teams to build controls that work in practice, not just on paper. You’ll execute technology audits, turn findings into practical improvements, and use data analytics, au

REMOTEgitrestagile
View job →
L
Lyft
📍 Mexico City• Full-time
23 days ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Every day, millions of riders and drivers depend on Lyft to get where they're going. When something goes wrong along the way, they expect us to make it right — quickly, clearly, and without friction. How fast and how well we resolve those moments shapes whether people keep choosing Lyft. The Self-Serve Intelligence team, within the Safety & Customer Care org, is composed of engineers building the AI-powered systems that do exactly that: resolving rider and driver suboptimal experiences without agent involvement through AI Assist (e.g. AI Agents), automations, and self-serve workflows. Our goal is to make getting help feel effortless. We design and build backend services, APIs, and GenAI-powered products that combine robust engineering with applied AI to deliver reliable, scalable self-serve experiences. We are looking for a highly motivated, collaborative, team-focused and technically strong Software Engineer to join our Self-Serve Intelligence team. As a member of this team, you will build the services and AI-powered products that resolve customer issues autonomously. Every day, you'll partner with machine learning engineers, product, design, data science, and operations on high-impact projects — from shipping new AI Agent capabilities, to building the evaluation pipelines that keep their quality high, to improving the backend services underneath them. You'll bring strong engineering instincts, genuine curiosity about applied AI, and a willingness to work through ambiguity in a space that changes month to month. Responsibilities: Write well-crafted, well-tested, readable, and maintainable code Partner with senior engineers to design, build, and ship backend services and GenAI-powered products (e.g. AI Agents) that resolve rider and driver suboptimal experiences Independently lead tasks fro

pythonjavaredis
View job →
O
OneTrust
📍 Atlanta• Full-time• From $116.5K/yr
15 days ago

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this role, you will part of the R&D Team that works on mission-critical applications. Your Mission Engage and partner with various Engineering, Operations, and Product teams to design, deliver, and maintain a highly available and performant application platform. Build and implement application observability and platform monitoring tools to continuously improve the customer experience Eliminate toil by automating processes, tuning alerts, and improving code where it is most needed Frequently evaluate new ideas and trends to identify potentially useful tools and techniques Collaborate with different functional groups to identify gaps, prioritize, and resolve issues Defining, implementing, and maintaining SLIs and SLOs aligned with customer experience. Design and instrument SLIs such as latency, error rates, and availability across critical services Manage and enforce error budgets to balance system reliability with product feature v

pythonjavasql
View job →
O
OneTrust
📍 Atlanta• Full-time• From $116.5K/yr
15 days ago

Strength in Trust OneTrust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, OneTrust is once again redefining what responsible innovation looks like. OneTrust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse at scale. Trusted by thousands of organizations worldwide, OneTrust is shaping the future where trusted data becomes a transformative force for business and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this role, you will part of the R&D Team that works on mission-critical applications. Your Mission Engage and partner with various Engineering, Operations, and Product teams to design, deliver, and maintain a highly available and performant application platform. Build and implement application observability and platform monitoring tools to continuously improve the customer experience Eliminate toil by automating processes, tuning alerts, and improving code where it is most needed Frequently evaluate new ideas and trends to identify potentially useful tools and techniques Collaborate with different functional groups to identify gaps, prioritize, and resolve issues Defining, implementing, and maintaining SLIs and SLOs aligned with customer experience. Design and instrument SLIs such as latency, error rates, and availability across critical services Manage and enforce error budgets to balance system reliability with product feature v

pythonjavasql
View job →
O
Okta
📍 Bengaluru• Full-time
18 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, we are building the future of secure, enterprise-grade cloud automation and system connectivity. We are looking for a Software Engineer to join our global Automation Engineering team to design, scale, and govern our enterprise integration substrate using AWS cloud services and modern iPaaS platforms. This is a individual contributor role for a hands-on system engineer who executes moderately complex tasks, builds platform components and collaborates under senior guidance. What You'll do : Contribute technical design and execution for automation initiatives within the team, creating paved paths that enable builders across Okta to connect enterprise systems seamlessly. Design, build, and deploy high-throughput event-driven integration flows , API gateways, and async orchestration workflows using AWS architectures and iPaaS platforms. Build reusable frameworks , developer SDKs, self-service primitives, and integration templates to streamline automation delivery. Develop and maintain AWS cloud services and modern iPaaS tooling, driving scalable architecture, reliability, and builder enablement. Partner with operations, security, and platform teams to strengthen monitoring, observability, structured audit logging, and automated governance. Actively participate in code reviews, team agile ceremonies, and technical discussions Collaborate with security, compliance, and business stakeholders to ensure self-service automations are secure, resilient, zero-tr

typescriptpythonjava
View job →
S
Stripe
📍 Wa Or New York• Full-time• Remote
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Finance and Strategy Data Science builds the forecasting models, data infrastructure, and analytics tools at the core of how Stripe measures and plans its business. The team owns everything from hierarchical time series and agentic forecasting tools that predict payment volumes and revenue margins, to the governed metrics platform that feeds company-wide dashboards and executive reporting. We partner closely with Finance and Strategy, GTM, and Product stakeholders to directly inform financial decisions across Stripe's entire business. The team combines technical depth, strategic thinking, and executive partnership that develops both technical and business expertise. What you’ll do Data Science Managers at Stripe are responsible for the success of their team. You'll be deeply involved in the modeling and design processes as well as coaching, mentoring, and leading the team. You'll have a deep understanding of how to drive efficient data science teams and you'll have a strong user-focus. You'll be working with data scientists, analysts, and engineers on creating technical solutions and communicating effectively across teams and senior leadership. Responsibilities Drive the roadmap and priorities for your team, and work with many Stripe leaders across the company to enhance our ability to be data-driven. Collaborate with stakeholders across the organization such as engineering, analytics, operations, finance, and marketing. Lead and manage proc

REMOTErestmachine learninggo
View job →
D
Discord
📍 San Francisco Bay Area• Full-time• $132K – $148.5K/yr
1mo ago

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly everyone does: play video games. Discord plays a uniquely important role in the future of gaming, and we are focused on making it easier and more fun for people to hang out before, during, and after playing games. Discord's mission is to give people the power to create space to find belonging in their lives, and every day millions of people trust Discord to be a place where they can talk, hang out, and build community. Keeping that space safe starts with teams like ours. Discord's Violent Actors Team, part of the broader Threat Operations org within Trust & Safety, focuses on identifying and disrupting violent and extremist networks and threat actors. As a Threat Investigator, you'll lead investigations into these communities, identify novel detection mechanisms, and develop enforcement strategies to stay ahead of ongoing and emerging threats. A typical day might include conducting deep-dive network investigations, compiling high-value signals to fuel proactive detection, or collaborating with internal and external partners on incident response. This role requires the ability to flex across a range of harm portfolios as team priorities evolve. This person will report to the team's Senior Manager. What You'll Be Doing Leading complex investigations into violent extremist and threat actor networks, developing and iterating on detection and disruption strategies across your investigative portfolio Executing on roadmap projects, including cross-functional coordination with policy, engineering, and ML/AI teams, documentation of behavioral trends, and identification of scalable mitigation strategies Leveraging investigative signals and subject matter expertise to inform threat prioritization decisions Identifying emergence of novel harm behaviors, and gaps in Discord's detection, policy, and enforcement capabilitie

restmachine learningai
View job →
O
1mo ago

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Safely delivering increasingly capable AI systems requires scalable technical safeguards, clear ownership of emerging risks, rigorous deployment readiness, and close coordination across research, engineering, product, operations, legal, policy, and external partners. Our Technical Program Managers lead complex, high-stakes initiatives that turn safety commitments into deployed systems and measurable outcomes. We work across model development, infrastructure, product, and operational response to help ensure our technology is deployed responsibly and cannot be used to cause serious real-world harm. About the Role We’re seeking Technical Program Managers to drive complex product, platform, and safety initiatives across ChatGPT, API, enterprise, and related deployment environments. These roles operate at the intersection of technical strategy and execution: you will turn safety and product priorities into actionable plans, influence architectural and operational decisions, and deliver durable capabilities across model, infrastructure, application, and platform layers. Depending on the role, you may enable sensitive or high-impact model deployments, integrate safeguards into cloud and API platforms, prevent violent misuse and other serious harms, improve detection and enforcement systems, create platform solutions for safety or establish new programs as risks evolve. You will partner deeply with engineers, researchers, product managers, and operational teams while communicating technical tradeoffs and program decisions to senior leadership. You bring technical fluency, product judgment, and a strong execution record. You’re comfortable navigating ambiguity, advocating for users and developers, balancing safety with model usefulness, and leading cross-functional work with urgency, rigor, and empathy. Specific focus areas and scope will vary by opening and level. Thi

awsrestmachine learning
View job →

About the Team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. Safely delivering increasingly capable AI systems requires scalable technical safeguards, clear ownership of emerging risks, rigorous deployment readiness, and close coordination across research, engineering, product, operations, legal, policy, and external partners. Our Technical Program Managers lead complex, high-stakes initiatives that turn safety commitments into deployed systems and measurable outcomes. We work across model development, infrastructure, product, and operational response to help ensure our technology is deployed responsibly and cannot be used to cause serious real-world harm. About the Role We’re seeking Technical Program Managers to drive complex product, platform, and safety initiatives across ChatGPT, API, enterprise, and related deployment environments. These roles operate at the intersection of technical strategy and execution: you will turn safety and product priorities into actionable plans, influence architectural and operational decisions, and deliver durable capabilities across model, infrastructure, application, and platform layers. Depending on the role, you may enable sensitive or high-impact model deployments, integrate safeguards into cloud and API platforms, prevent violent misuse and other serious harms, improve detection and enforcement systems, create platform solutions for safety or establish new programs as risks evolve. You will partner deeply with engineers, researchers, product managers, and operational teams while communicating technical tradeoffs and program decisions to senior leadership. You bring technical fluency, product judgment, and a strong execution record. You’re comfortable navigating ambiguity, advocating for users and developers, balancing safety with model usefulness, and leading cross-functional work with urgency, rigor, and empathy. Specific focus areas and scope will vary by opening and level. In

awsrestmachine learning
View job →

SUMMARY STATEMENT We are looking for a Solution Architect to design the technical solutions behind our client engagements and give delivery teams a clear, workable path from concept to production. You will work across enterprise data, software applications and GenAI - translating complex business problems into practical architectures that delivery teams can build and scale. This could include architecting an agentic workflow for clinical operations, a conversational analytics product grounded in enterprise data, or an AI-enabled decision platform for commercial teams. You will work directly with clients, define the architecture, test the most important technical decisions yourself and establish the foundations for successful delivery. This is an architecture-first role with meaningful hands-on engineering: you will stay close enough to implementation to prove the architecture works and support it through production delivery, without becoming the primary engineer for every component. You will also help shape the reusable patterns, technical standards and accelerators behind Lynx’s growing AI-native life sciences practice. KEY RESPONSIBILITIES Solution Architecture Own the end-to-end solution architecture for client engagements, including data models, system design, integration patterns and technology choices. Translate business requirements into clear technical designs and implementation paths that delivery teams can build from. Design solutions spanning enterprise data, APIs, applications, cloud platforms and GenAI capabilities. Lead technical discovery with clients: understand requirements, assess existing systems and identify dependencies, constraints and delivery risks. Present architectural options and trade-offs clearly to technical teams, business stakeholders and senior leaders. Make pragmatic decisions across build speed, cost, scalability, security and maintainability. Review key implementation decisions and remain

typescriptpythonaws
View job →
🔔

Get new senior machine learning operations engineer jobs by email

Daily job updates · Unsubscribe anytime