Jobiba hiring network

Infrastructure And Mlops Engineer Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current infrastructure and mlops engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

G
greenhouse,Graphcore
📍 CambridgeFull-time
3 hrs ago

About Graphcore At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence . Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support AI research and engineering teams Deploy and maintain services with Kubernetes and Docker Manage our Cloud Infrastructure using tools such as Terraform Candidate Profile Essential:

pythonjavaaws
View job →
W
Wellhub
📍 BrazilFull-timeRemote
1 hr ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Senior MLOps Engineer to our Product Development team in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. Join the ML Development Lifecycle team within our Product Development (PD) organisation, where we are redefining how a global tech company leverages intelligence. We build the foundations that allow hundreds of engineers and data scientists to develop and deploy AI at scale. You will own the evolution of our cloud-native ecosystem , creating a seamless and high-performance environment for the next generation of AI-driven products. If you are a software-minded engineer who thrives at the intersection of scalable Infrastructure and ML/AI orchestration, this is your chance to build a world-class platform that serves millions of users worldwide. YOUR IMPACT Scale the Ecosystem: Evolve and maintain our Kubeflow, Feast and Spark-on-Kubernetes infrastructure, ensuring it can handle the increasing complexity of both traditional ML and the new wave of AI. Build for Autonomy:

REMOTEpythonawskubernetes
View job →
J
Jumio
📍 IndiaFull-timeRemote
1 hr ago

Role Purpose: At Jumio, you will work for one of the market leaders in the global identity verification space that is helping to make the digital world a safer place for everyone. As a Software Development Engineer in the MLOpsTeam, you will develop the blueprint for highly scalable and performant ML model serving. Role Value: As a Software Engineer (SDE III), you will drive the continuous improvement of the infrastructure and applications to manage the lifecycle of ML assets (data, models) to better developer experience and strengthen governance capabilities. Secondly, you will design and implement robust ML infrastructure for model deployment, serving, and optimization. You will work on efficient CI/CD pipelines for ML models and leverage advanced compilers or hardware optimization to maximize inference performance while optimizing costs. We welcome you to challenge us to impact our software development processes and tools. Example Responsibilities: Upgrade ML assets (models, data) management systems for better developer experience and robust governance capabilities Build and optimize model serving infrastructure with a focus on inference latency and cost optimization Architect efficient inference pipelines that balance latency, throughput, and cost across various acceleration options Implement cost-efficient, enterprise-scale solutions Collaborate in a cross-functional, distributed team for continuous system improvement Work with MLEs, QA Engineers, and DevOps Engineers Evaluate and implement new technologies and tools Contribute to architectural decisions for distributed ML systems Experience and Qualifications : 5+ years of experience in software engineering with Python Experience with model lifecycle management (MLFlow, Weights & Biases or equivalent) Experience with data management ecosystem (quality, transformation, catalog) Experience with ML frameworks, particularly PyTorch Experience optimizing ML models with hardwar

REMOTEpythonawsdocker
View job →
E
Everpure
📍 BengaluruFull-time
1 hr ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pure Solutions team as a Senior MLOps Solutions Engineer to architect and build high-scale, enterprise-grade AI/ML solutions. You will be instrumental in integrating Pure Storage platforms with the evolving open-source MLOps ecosystem (Kubeflow, MLflow, Ray) to operationalize the complete machine learning lifecycle. This role requires a creative technologist with deep Python expertise to drive innovation and enable our customers and partners to achieve production AI success. WHAT YOU'LL DO Design and Automate MLOps Pipelines: Lead the development of end-to-end MLOps workflows using CI/CD tools (Git/Jenkins) and orchestration platforms (MLflow/Kubeflow), specifically integrating Pure Storage's FlashBlade, FlashArray, and Portworx as the high-performance data plane for data ingestion, training, and inference. Build High-Performance AI/ML Reference Architectures: Create validated, repeatable deployment models using Infrastructure as Code (e.g., Ansible, Terraform) for AI/ML environments spanning bare metal, virtual machines, and GPU-accelerated Kubernetes clusters, ensuring optimal performance for distributed training. Optimize and Operationalize GPU Inference: Architect and implement solutions for high-throughput, low-latency model serving, utilizing technologies like NVIDIA Triton Inference Server and advanced optimization techniques (quantization, model sharding like DeepSpeed/Megatron-LM, and dynamic bat

pythonawskubernetes
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe processes over $1T in payments volume per year, which is roughly 1% of the world’s GDP. The tremendous amount of data makes Stripe one of the best places to do machine learning. The ML Infra team builds services and tools that power every step in the ML lifecycle, including data exploration, feature generation, experimentation, training, deploying, serving ML models, and building LLM applications. With the phenomenal developments happening in the field of AI, we are positioned to accelerate the adoption of AI/ML across all parts of the company by building highly scalable and reliable foundational infrastructure. What you’ll do You will work closely with machine learning engineers, data scientists, and product engineering teams to enable seamless end-to-end experience in building solutions across data, analytics, and AI/ML platforms. You will build the next generation of ML Infra services and major new capabilities that substantially improve ML development velocity and MLOps maturity across the company. Responsibilities Designing and building scalable, reliable, and secure services for notebooks, ML model training, experimentation, serving, and LLM applications across multiple regions. Creating services and libraries that enable ML engineers at Stripe to seamlessly transition from experimentation to production across Stripe’s systems. Working directly with product teams and ML engineers to improve their day-to-day pr

restmachine learningai
View job →

AI/ML Dev - Chatbots • 8+ years of experience in Data/AI Projects • Understanding of end-to-end architecture for Generative AI solutions aligned with business goals. • Experience in Azure OpenAI integration (GPT models, embeddings) with prompt engineering and model tuning. • Programming experience in Python for AI project is a must • Designs scalable RAG systems using Azure AI Search, vector databases, and secure data pipelines. • Knowledge of MLOps and CI/CD workflows using Azure DevOps and automated testing frameworks. • Establishes Python coding standards, reviews code, and mentors development teams. • Knowledge of deployment and governance of AI applications across Azure infrastructure. • Work with cross function teams (IT/ Non IT) to help the development teams build the solutions faster and more efficiently.

G
greenhouse,Guidepoint
📍 TorontoFull-timeC$135K – C$210K/yr
4 hrs ago

Overview: Guidepoint seeks an experienced Data/AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The Senior AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization

pythonreactaws
View job →
G
greenhouse,Guidepoint
📍 TorontoFull-timeC$135K – C$210K/yr
4 hrs ago

Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-based AI team. The Toronto Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible AI. This strategic investment is integral to Guidepoint’s vision for the future, aiming to develop cutting-edge Generative AI and analytical capabilities that will underpin Guidepoint’s Next-Gen research enablement platform and data products. This role demands exceptional leadership and technical prowess to drive the development of next-generation research enablement platforms and AI-driven data products. You will develop and scale Generative AI-powered systems, including large language model (LLM) applications and research agents, while ensuring the integration of responsible AI and best-in-class MLOps. The AI/ML Engineer will be a primary contributor to building scalable AI/ML capabilities using Databricks and other state-of-the-art tools across all of Guidepoint’s products. Guidepoint’s Technology team thrives on problem-solving and creating happier users. As Guidepoint works to achieve its mission of making individuals, businesses, and the world smarter through personalized knowledge-sharing solutions, the engineering team is taking on challenges to improve our internal application architecture and create new AI-enabled products to optimize the seamless delivery of our services. This is a hybrid position based in Toronto. What You'll Do: Architect and Build Production Systems: Design, build, and operate scalable, low-latency backend services and APIs that serve Generative AI features, from retrieval-augmented generation (RAG) pipelines to complex agentic systems. Own the AI Application Lifecycle: Own the end-to-end lifecycle of AI-powered applications, including system design, development, deployment (CI/CD), monitoring, and optimization in production

pythonreactaws
View job →

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Smartsheet is hiring a Senior Machine Learning Operations Engineer to architect our machine learning production lifecycle. Your mission is to maintain and deploy ML models to a scalable, reliable, and secure production environment. You will design and maintain the infrastructure, automation, and monitoring systems that ensure our AI products are high-performing and cost-effective. You will report to our Director, Analytics Engineering & Data Governance and work from our Bangalore, India office. You Will: Model and Pipeline Automation Automate the deployment and retraining of ML models, from training through to production inference, by building and managing complete CI/CD/CT (Continuous Training) pipelines, adhering to MLOps best practices. Build, fine-tune, or use pre-trained LLMs, deep learning models or traditional machine learning models. Evaluate and recommend AI or ML solutions for the product using any combination of vendor solutions and/or custom-built models. Governance & Compliance Implement model versioning, lineage tracking, and auditing to ensure compliance with security and ethical standards. Performance Monitoring Continuously monitor the health and performance of production machine learning models, proactively identifying and correcting model drift, staleness, and performance degradation. Incorporate user feedback for iterative improvements and manage necessary model retraining cycles. Cross-Functional Collaboration Act as the "glue" between Data Scientists (who build models

pythonawsazure
View job →
J
1 hr ago

Backend Engineer (Senior Level) - SDE IV We're looking for a Senior Backend Engineer to lead the architecture and evolution of backend services that deploy and serve machine learning models in production. You'll work closely with ML Engineers, Platform, and Product teams to build scalable, reliable systems and drive technical direction across multiple teams. What You’ll Do Design and drive the long-term architecture of backend services for biometrics and ML model serving. Collaborate with core platform and backend teams on organization-wide architectural initiatives. Partner with business and engineering teams to design and deliver cross-cutting platform capabilities. Lead architectural reviews, mentor engineers, and promote engineering best practices. Build and maintain backend services for deploying and serving ML models Monitor service reliability, performance, and scalability in production Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 + Kubernetes Support real-time and batch inference workflows Contribute to CI/CD pipelines and deployment automation What We’re Looking For Strong expertise in backend development using Java and working knowledge of Python. Experience mentoring engineers and driving architectural decisions. Working knowledge of Python, especially for ML-related workflows Hands-on experience with AWS (e.g., DynamoDB, ECS, EC2, Redis, S3, SageMaker) Familiarity with Terraform or other infrastructure-as-code tools, and experience with CI/CD and production monitoring Experience with observability tools (Datadog, New Relic, etc.) Experience with containers and orchestration (Docker, ECS, etc.) Understanding of how ML models are deployed and served in production Experience with Kubernetes Nice to Have Experience with MLOps or ML platform engineering. Experience with asynchronous programming and event-driven systems. Jumio Values: IDEAL: Integrity, Diversity, Empowerment, Accountability, Leading Innovation Equal Opportunities :

REMOTEpythonjavaredis
View job →
D
DataCamp
📍 BelgiumFull-time
1 hr ago

Data and AI skills are critical for thriving today, and DataCamp is the platform that empowers everyone to learn them. We help individuals and Fortune 1000 companies close the data and AI skills gap through world-class learning, hands-on training, and a global community of expert instructors. In this role, you'll combine deep cloud expertise, technical breadth, and a passion for great didactic experiences to build the next generation of cloud education at DataCamp . About the role This is an individual contributor role. You will collaborate with subject matter experts and leverage in-house-built cutting-edge AI tooling to scale high-quality course creation across cloud platforms (AWS, Azure, GCP) and cloud-adjacent topics such as DevOps, infrastructure-as-code, MLOps, and cloud data engineering. Here's what your day-to-day will look like: Manage the entire content development lifecycle and deadlines. Source and recruit top-tier subject-matter experts as instructors. Collaborate with instructors to create engaging content. Consistently leverage AI tools like Claude Code, Cursor, and more to drive high-quality content production at scale. Design, review, and create content on Cloud platforms, cloud architecture, DevOps, cloud data engineering, and related topics. You will review content from a learner perspective, ensuring it is technically accurate and pedagogically effective. Continuously assess course performance using learner feedback and engagement data to drive improvements. Identify and prioritize existing curriculum gaps or new topics in DataCamp's cloud curriculum. We’re excited about you because you have the following A solid hands-on background working with one or more major cloud platforms (AWS, Azure, or GCP). Relevant cloud certifications (e.g. AWS Solutions Architect, Google Professional Cloud Architect, Azure Administrator) are a strong plus. Broad cloud fluency spanning infrastructure, storage, compute, networking, and managed services, with additiona

pythonsqlaws
View job →
D
DataCamp
📍 BelgiumFull-time
1 hr ago

Data and AI skills are critical for thriving today, and DataCamp is the platform that empowers everyone to learn them. We help individuals and Fortune 1000 companies close the data and AI skills gap through world-class learning, hands-on training, and a global community of expert instructors. In this role, you'll combine deep cloud expertise, technical breadth, and a passion for great didactic experiences to build the next generation of cloud education at DataCamp . About the role This is an individual contributor role. You will collaborate with subject matter experts and leverage in-house-built cutting-edge AI tooling to scale high-quality course creation across cloud platforms (AWS, Azure, GCP) and cloud-adjacent topics such as DevOps, infrastructure-as-code, MLOps, and cloud data engineering. Here's what your day-to-day will look like: Manage the entire content development lifecycle and deadlines. Source and recruit top-tier subject-matter experts as instructors. Collaborate with instructors to create engaging content. Consistently leverage AI tools like Claude Code, Cursor, and more to drive high-quality content production at scale. Design, review, and create content on Cloud platforms, cloud architecture, DevOps, cloud data engineering, and related topics. You will review content from a learner perspective, ensuring it is technically accurate and pedagogically effective. Continuously assess course performance using learner feedback and engagement data to drive improvements. Identify and prioritize existing curriculum gaps or new topics in DataCamp's cloud curriculum. We’re excited about you because you have the following A solid hands-on background working with one or more major cloud platforms (AWS, Azure, or GCP). Relevant cloud certifications (e.g. AWS Solutions Architect, Google Professional Cloud Architect, Azure Administrator) are a strong plus. Broad cloud fluency spanning infrastructure, storage, compute, networking, and managed services, with additiona

pythonsqlaws
View job →
D
DataCamp
📍 BelgiumFull-time
1 hr ago

Data and AI skills are critical for thriving today, and DataCamp is the platform that empowers everyone to learn them. We help individuals and Fortune 1000 companies close the data and AI skills gap through world-class learning, hands-on training, and a global community of expert instructors. In this role, you'll combine deep cloud expertise, technical breadth, and a passion for great didactic experiences to build the next generation of cloud education at DataCamp . About the role This is an individual contributor role. You will collaborate with subject matter experts and leverage in-house-built cutting-edge AI tooling to scale high-quality course creation across cloud platforms (AWS, Azure, GCP) and cloud-adjacent topics such as DevOps, infrastructure-as-code, MLOps, and cloud data engineering. Here's what your day-to-day will look like: Manage the entire content development lifecycle and deadlines. Source and recruit top-tier subject-matter experts as instructors. Collaborate with instructors to create engaging content. Consistently leverage AI tools like Claude Code, Cursor, and more to drive high-quality content production at scale. Design, review, and create content on Cloud platforms, cloud architecture, DevOps, cloud data engineering, and related topics. You will review content from a learner perspective, ensuring it is technically accurate and pedagogically effective. Continuously assess course performance using learner feedback and engagement data to drive improvements. Identify and prioritize existing curriculum gaps or new topics in DataCamp's cloud curriculum. We’re excited about you because you have the following A solid hands-on background working with one or more major cloud platforms (AWS, Azure, or GCP). Relevant cloud certifications (e.g. AWS Solutions Architect, Google Professional Cloud Architect, Azure Administrator) are a strong plus. Broad cloud fluency spanning infrastructure, storage, compute, networking, and managed services, with additiona

pythonsqlaws
View job →
D
DataCamp
📍 BelgiumFull-time
1 hr ago

Data and AI skills are critical for thriving today, and DataCamp is the platform that empowers everyone to learn them. We help individuals and Fortune 1000 companies close the data and AI skills gap through world-class learning, hands-on training, and a global community of expert instructors. In this role, you'll combine deep cloud expertise, technical breadth, and a passion for great didactic experiences to build the next generation of cloud education at DataCamp . About the role This is an individual contributor role. You will collaborate with subject matter experts and leverage in-house-built cutting-edge AI tooling to scale high-quality course creation across cloud platforms (AWS, Azure, GCP) and cloud-adjacent topics such as DevOps, infrastructure-as-code, MLOps, and cloud data engineering. Here's what your day-to-day will look like: Manage the entire content development lifecycle and deadlines. Source and recruit top-tier subject-matter experts as instructors. Collaborate with instructors to create engaging content. Consistently leverage AI tools like Claude Code, Cursor, and more to drive high-quality content production at scale. Design, review, and create content on Cloud platforms, cloud architecture, DevOps, cloud data engineering, and related topics. You will review content from a learner perspective, ensuring it is technically accurate and pedagogically effective. Continuously assess course performance using learner feedback and engagement data to drive improvements. Identify and prioritize existing curriculum gaps or new topics in DataCamp's cloud curriculum. We’re excited about you because you have the following A solid hands-on background working with one or more major cloud platforms (AWS, Azure, or GCP). Relevant cloud certifications (e.g. AWS Solutions Architect, Google Professional Cloud Architect, Azure Administrator) are a strong plus. Broad cloud fluency spanning infrastructure, storage, compute, networking, and managed services, with additiona

pythonsqlaws
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform. You Will: Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with

pythonsqlaws
View job →
🔔

Get new infrastructure and mlops engineer jobs by email

Daily job updates · Unsubscribe anytime