Jobiba hiring network

Senior Mlops Engineer Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current senior mlops engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

W
Wellhub
📍 BrazilFull-timeRemote
3 days ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Senior MLOps Engineer to our Product Development team in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. Join the ML Development Lifecycle team within our Product Development (PD) organisation, where we are redefining how a global tech company leverages intelligence. We build the foundations that allow hundreds of engineers and data scientists to develop and deploy AI at scale. You will own the evolution of our cloud-native ecosystem , creating a seamless and high-performance environment for the next generation of AI-driven products. If you are a software-minded engineer who thrives at the intersection of scalable Infrastructure and ML/AI orchestration, this is your chance to build a world-class platform that serves millions of users worldwide. YOUR IMPACT Scale the Ecosystem: Evolve and maintain our Kubeflow, Feast and Spark-on-Kubernetes infrastructure, ensuring it can handle the increasing complexity of both traditional ML and the new wave of AI. Build for Autonomy:

REMOTEpythonawskubernetes
View job →
H
27 days ago

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
E
Everpure
📍 BengaluruFull-time
3 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat

pythonawskubernetes
View job →
G
Guidepoint
📍 TorontoFull-timeC$135K – C$210K/yr
3 days ago

Overview: Guidepoint seeks an experienced Data/AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The Senior AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization

pythonreactaws
View job →
DC
3 days ago

Must be based in Vancouver The role We're hiring a dedicated data engineer to own the production data platform that our delivery, product, and engineering teams run on; designing integrated, governed data pipelines and delivering automated reporting, AI-assisted workflows, and predictive signals on top of them. You'll write production code, design systems, own CI/CD, and be accountable for the correctness of data that leaders make decisions on. What you'll do Design and operate our cloud data platform: ingestion, transformation, orchestration and serving. Integrate data from across the business (delivery tooling, CRM, product telemetry, finance, support and customer feedback systems) with shared identifiers, data contracts and lineage. Build automated and continuously refreshed reporting so teams manage by exception rather than chasing status. Connect approved AI agents to governed data with structured outputs, provenance, guardrails and human approval in the loop. Build feature pipelines and the MLOps controls behind predictive use cases: tests, versioning, promotion gates and drift monitoring. Own the engineering standards for data: testing, observability, environment promotion, PII classification and access control. What you'll bring Strong software engineering fundamentals: production-quality code, API and interface design, testing discipline, systems design. Real experience building and operating production data platforms on a cloud warehouse or lakehouse (Snowflake and AWS preferred) with dbt and a modern orchestrator. Practical AI tooling experience: something shipped, not prototyped. LLM-backed classification, extraction or structured-output pipelines; agent and tool-calling workflows; retrieval; evals. You can reaso

awsci/cdgit
View job →

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is hiring a Senior Machine Learning Operations Engineer to architect our machine learning production lifecycle. Your mission is to maintain and deploy ML models to a scalable, reliable, and secure production environment. You will design and maintain the infrastructure, automation, and monitoring systems that ensure our AI products are high-performing and cost-effective. You will report to our Director, Analytics Engineering & Data Governance and work from our Bangalore, India office. You Will: Model and Pipeline Automation Automate the deployment and retraining of ML models, from training through to production inference, by building and managing complete CI/CD/CT (Continuous Training) pipelines, adhering to MLOps best practices. Build, fine-tune, or use pre-trained LLMs, deep learning models or traditional machine learning models. Evaluate and recommend AI or ML solutions for the product using any combination of vendor solutions and/or custom-built models. Governance & Compliance Implement model versioning, lineage tracking, and auditing to ensure compliance with security and ethical standards. Performance Monitoring Continuously monitor the health and performance of production machine learning models, proactively identifying and correcting model drift, staleness, and performance degradation. Incorporate user feedback for iterative improvements and manage necessary model retraining cycles. Cross-Functional Collaboration Act as the "glue" between Data Scientists (who build models

pythonawsazure
View job →

About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries. What You’ll Do: Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference. Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang . You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience. Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence. Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid. What are we looking for Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructu

pythonkubernetesci/cd
View job →
J
3 days ago

Backend Engineer (Senior Level) - SDE IV We're looking for a Senior Backend Engineer to lead the architecture and evolution of backend services that deploy and serve machine learning models in production. You'll work closely with ML Engineers, Platform, and Product teams to build scalable, reliable systems and drive technical direction across multiple teams. What You’ll Do Design and drive the long-term architecture of backend services for biometrics and ML model serving. Collaborate with core platform and backend teams on organization-wide architectural initiatives. Partner with business and engineering teams to design and deliver cross-cutting platform capabilities. Lead architectural reviews, mentor engineers, and promote engineering best practices. Build and maintain backend services for deploying and serving ML models Monitor service reliability, performance, and scalability in production Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 + Kubernetes Support real-time and batch inference workflows Contribute to CI/CD pipelines and deployment automation What We’re Looking For Strong expertise in backend development using Java and working knowledge of Python. Experience mentoring engineers and driving architectural decisions. Working knowledge of Python, especially for ML-related workflows Hands-on experience with AWS (e.g., DynamoDB, ECS, EC2, Redis, S3, SageMaker) Familiarity with Terraform or other infrastructure-as-code tools, and experience with CI/CD and production monitoring Experience with observability tools (Datadog, New Relic, etc.) Experience with containers and orchestration (Docker, ECS, etc.) Understanding of how ML models are deployed and served in production Experience with Kubernetes Nice to Have Experience with MLOps or ML platform engineering. Experience with asynchronous programming and event-driven systems. Jumio Values: IDEAL: Integrity, Diversity, Empowerment, Accountability, Leading Innovation Equal Opportunities :

REMOTEpythonjavaredis
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform. You Will: Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with

pythonsqlaws
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform. You Will: Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with

pythonsqlaws
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Technology Eva

pythonsqlaws
View job →
T
20 days ago

At Trustpilot, we're on an incredible journey. We're a profitable, high-growth FTSE-250 company with a big vision: to become the universal symbol of trust. We run the world's largest open customer review platform, and while we've come a long way, there's still so much exciting work to do. Come join us at the heart of trust! We are growing our engineering team at Trustpilot and are looking to welcome a Software Engineer I into the Trust Tech department! You will join a cross-functional team with full ownership of our products and codebase, where you will take part in every step of the development process, from ideation to maintenance. This is a place where you can grow as an individual, learn from senior mentors, and have a real influence on the direction of our projects. About the team: You will be joining a brand new team being built within the Trust Tech department, focusing specifically on businesses. The team is responsible for building and maintaining automated systems that enforce Trustpilot’s terms of use for businesses. This includes developing systems for automatic detection, enforcement of misuse, and internal investigation tools. This work is vital to maintaining Trust and Transparency on the platform. To achieve this, we collaborate closely with Data Scientists, Data Ops, and ML Ops to build sophisticated automated architectures that integrate detection models and robustly scale our misuse enforcement. What you’ll be doing: Work in a cross-functional “full ownership” team alongside Product, Design, and Data Science. Implement and release new features with a focus on backend stability and scalability. Help build solutions to handle high-volume data processing for fraud detection. Maintain and improve the internal tooling frontends (React) used for investigations. Troubleshoot existing software, squash bugs, and learn how to optimize database performance. Participate in technical discussions and learn best practices for Infrastructure a

javascripttypescriptjava
View job →
T
20 days ago

At Trustpilot, we're on an incredible journey. We're a profitable, high-growth FTSE-250 company with a big vision: to become the universal symbol of trust. We run the world's largest open customer review platform, and while we've come a long way, there's still so much exciting work to do. Come join us at the heart of trust! We are growing our engineering team at Trustpilot and are looking to welcome a Software Engineer I into the Trust Tech department! You will join a cross-functional team with full ownership of our products and codebase, where you will take part in every step of the development process, from ideation to maintenance. This is a place where you can grow as an individual, learn from senior mentors, and have a real influence on the direction of our projects. About the team: You will be joining a brand new team being built within the Trust Tech department, focusing specifically on businesses. The team is responsible for building and maintaining automated systems that enforce Trustpilot’s terms of use for businesses. This includes developing systems for automatic detection, enforcement of misuse, and internal investigation tools. This work is vital to maintaining Trust and Transparency on the platform. To achieve this, we collaborate closely with Data Scientists, Data Ops, and ML Ops to build sophisticated automated architectures that integrate detection models and robustly scale our misuse enforcement. What you’ll be doing: Work in a cross-functional “full ownership” team alongside Product, Design, and Data Science. Implement and release new features with a focus on backend stability and scalability. Help build solutions to handle high-volume data processing for fraud detection. Maintain and improve the internal tooling frontends (React) used for investigations. Troubleshoot existing software, squash bugs, and learn how to optimize database performance. Participate in technical discussions and learn best practices for Infrastructure a

javascripttypescriptjava
View job →
T
Trustpilot
📍 CopenhagenFull-time
20 days ago

At Trustpilot, we're on an incredible journey. We're a profitable, high-growth FTSE-250 company with a big vision: to become the universal symbol of trust. We run the world's largest open customer review platform, and while we've come a long way, there's still so much exciting work to do. Come join us at the heart of trust! We are growing our engineering team at Trustpilot and are looking to welcome a Software Engineer I into the Trust Tech department! You will join a cross-functional team with full ownership of our products and codebase, where you will take part in every step of the development process, from ideation to maintenance. This is a place where you can grow as an individual, learn from senior mentors, and have a real influence on the direction of our projects. About the team: You will be joining a brand new team being built within the Trust Tech department, focusing specifically on businesses. The team is responsible for building and maintaining automated systems that enforce Trustpilot’s terms of use for businesses. This includes developing systems for automatic detection, enforcement of misuse, and internal investigation tools. This work is vital to maintaining Trust and Transparency on the platform. To achieve this, we collaborate closely with Data Scientists, Data Ops, and ML Ops to build sophisticated automated architectures that integrate detection models and robustly scale our misuse enforcement. What you’ll be doing: Work in a cross-functional “full ownership” team alongside Product, Design, and Data Science. Implement and release new features with a focus on backend stability and scalability. Help build solutions to handle high-volume data processing for fraud detection. Maintain and improve the internal tooling frontends (React) used for investigations. Troubleshoot existing software, squash bugs, and learn how to optimize database performance. Participate in technical discussions and learn best practices for Infrastructure a

javascripttypescriptjava
View job →
I
4 hrs ago

Job Details: Job Description: The Role and Impact As a Systems and Solutions Engineer, you will drive the design, development, and integration of systems that combine software, firmware, board, and silicon/SoC components to meet specific customer needs. In this role, you will play a key part in defining, implementing, and optimizing solutions to ensure high performance, reliability, and quality across the system lifecycle. Your work will directly impact the seamless functionality and user experience of cutting-edge technologies, enhancing Intel's position in delivering innovative systems to global customers. Business Group You will be joining the Silicon and Platform Engineering Group (SPE), an organization committed to advancing Intel's mission of delivering world-class silicon and platform solutions. The group focuses on developing integrated systems that align with customer needs and support Intel's broader goals of leadership in technology innovation. SPE collaborates across diverse domains to ensure Intel platforms meet performance, reliability, and scalability requirements. Key Responsibilities - Design and develop software, firmware, and hardware solutions that integrate seamlessly across system components. - Lead the definition and implementation of system architecture, translating business opportunities into technical requirements and use cases. - Evaluate technical risk and optimize systems for ease of use, reliability, security, availability, and sustainability. - Drive technical solutions to address customer challenges, deploying systems and conducting benchmarks to validate performance. - Collaborate with cross-functional teams to influence next-generation requirements and solutions, guiding research and academic collaborations as needed. - Conduct lab experiments to simulate real-life environments, analyze prototype performance, and refine system

recruitment
View job →
🔔

Get new senior mlops engineer jobs by email

Daily job updates · Unsubscribe anytime