Jobiba hiring network

Software Engineer Ml Infrastructure Platform Salary India Jobs

6,326 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current software engineer ml infrastructure platform salary india jobs. Use filters to narrow by work mode, employment type, experience and date posted.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Product Security team ensures that Snowflake products are built and shipped with the highest level of security. Our team drives the security posture of Snowflake products and is responsible for embedding security into every stage of the product lifecycle, from design through deployment and beyond. We design and build frameworks, systems and services that keep Snowflake secure. As a Principal Software Engineer II on the Product Security team, you will be the senior technical authority for Product Security and play a critical leadership role in shaping and advancing Snowflake’s security. This is a unique opportunity to define and influence our long-term security strategy and have a direct impact on the security of the Snowflake platform and the trust of our customers. You will operate across organizational boundaries, guiding major security initiatives, influencing architectural decisions at the highest levels, setting the technical direction for the organization, and ensuring consistent security excellence across all product teams while working closely with business leaders to advance Snowflake’s business. The role requires deep expertise in security, software engineering, distributed systems, software infrastructure, AI/ML, applied cryptography, threat modeling and clou

pythonjavaai
View job →
W
Wellhub
📍 Brazil• Full-time• Remote
17 days ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Senior AI Platform Engineer to our Product Development team in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. Join the ML Development Lifecycle team within our Product Development (PD) organisation, where we are redefining how a global tech company leverages intelligence. We build the foundations that allow hundreds of engineers and data scientists to develop and deploy AI at scale. You will own the evolution of our cloud-native ecosystem , creating a seamless and high-performance environment for the next generation of AI-driven products. If you are a software-minded engineer who thrives at the intersection of scalable Infrastructure and ML/AI orchestration, this is your chance to build a world-class platform that serves millions of users worldwide. YOUR IMPACT Scale the Ecosystem: Evolve and maintain our Kubeflow, Feast and Spark-on-Kubernetes infrastructure, ensuring it can handle the increasing complexity of both traditional ML and the new wave of AI. Build for Autonomy:

REMOTEpythonawskubernetes
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Capital provides access to fast, flexible financing to small-and-medium businesses on Stripe to accelerate their growth, and we lent over $1B in 2024. Businesses use the funds for marketing, team growth, geographic expansion, working capital, new equipment purchases, and much more. Machine learning is core to Stripe Capital’s business—we use information about businesses from their activity within and outside of Stripe and our models to automatically underwrite uniquely tailored financing offers to their needs, which banks are often unable to do. We are doing so through models with an established performance history, data infrastructure that is Stripe scale, and a strong feedback loop that includes explainability, anomaly detection and a risk portfolio management layer. We're an end-to-end team going from ideas to models to shipping in production. What you’ll do As a machine learning engineer for Stripe Capital, you'll be responsible for designing, building, training, evaluating, deploying, and owning ML models in production with the goals of providing financing opportunities to as many users as possible while satisfying financial performance goals. You'll work closely with software engineers, data scientists, product managers, and risk managers to operate Stripe’s ML powered systems, features, and products. You'll also contribute to and influence ML architecture at Stripe and be a part of a larger ML community. Responsibilities Design

machine learningaigo
View job →
B
17 days ago

At Breeze, we're building the AI-powered infrastructure layer for global commerce, making it radically simpler for businesses to sell, get paid, and operate across markets. We go far beyond traditional payment processing. Breeze combines global payments, AI, stablecoins, and a Merchant of Record-like model to take on the complexity businesses typically manage themselves, including compliance, risk, fraud, chargebacks, reconciliation, and customer support. Our goal is simple: let businesses focus on building and selling great products while Breeze handles the complexity behind getting paid. Backed by Sequoia Capital , Multicoin Capital , and The Chainsmokers , Breeze is a successful, rapidly growing, and exceptionally well-capitalized company. We have the runway to think long term while remaining early enough that every person joining today can have a meaningful impact on what we build. We are hiring a Staff Machine Learning Engineer, Risk! As our Staff Machine Learning Engineer, Risk, you'll lead the evolution of our ML platform for payment risk, building the production-grade capabilities behind feature engineering, model training, deployment, monitoring, and continuous improvement. Risk decisions sit at the center of our business, and you'll own how those models get built, shipped, and kept healthy. This role reports to the CTO. You'll work closely with Risk, Software Engineering, and Data Engineering, and you'll be the senior technical voice for ML on the risk team. We're looking for someone who thrives in fast-moving environments, wants meaningful ownership, and is excited to build rather than simply maintain. What You'll Do Design and build ML infrastructure for payment risk detection, using Databricks as the core platform, in close partnership with software and data engineers. Bring structure to the team's ML environment: feature pipelines, versioning, job orchestration, and monitoring. Design and productionize models rather than just prototype them, including

machine learningaigo
View job →
W
Wellhub
📍 Brazil• Full-time• Remote
17 days ago

Your wellbeing, our mission. Join a company shaping a healthier world. GET TO KNOW US At Wellhub we're revolutionizing workplace wellness. Our platform connects employees worldwide to the best partners for fitness, mindfulness, therapy, nutrition, and sleep—all in one simple subscription. Headquartered in NYC with team members in Europe, North America and South America, we’re on a mission to make every company a wellness company. We believe work should be fulfilling, inspiring, and balanced. Here, you’ll find a team that values wellbeing, collaboration, and different perspectives, where passion and creativity push boundaries to create real impact. Your contributions will help shape a healthier, more balanced world for you and millions of people globally. Join us in redefining the future of wellbeing! THE OPPORTUNITY We are hiring a Senior MLOps Engineer to our Product Development team in Brazil ! This is a Remote – Brazil position, meaning you can work from anywhere within the country. Please note that this role is only open to candidates in Brazil. Join the ML Development Lifecycle team within our Product Development (PD) organisation, where we are redefining how a global tech company leverages intelligence. We build the foundations that allow hundreds of engineers and data scientists to develop and deploy AI at scale. You will own the evolution of our cloud-native ecosystem , creating a seamless and high-performance environment for the next generation of AI-driven products. If you are a software-minded engineer who thrives at the intersection of scalable Infrastructure and ML/AI orchestration, this is your chance to build a world-class platform that serves millions of users worldwide. YOUR IMPACT Scale the Ecosystem: Evolve and maintain our Kubeflow, Feast and Spark-on-Kubernetes infrastructure, ensuring it can handle the increasing complexity of both traditional ML and the new wave of AI. Build for Autonomy:

REMOTEpythonawskubernetes
View job →
O
Okta
📍 San Francisco• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Identity Threat Protection Team Identity Threat Protection (ITP), powered by cutting-edge Okta AI, shatters the mold of traditional login security. We deliver real-time, surgical defense against the most sophisticated and rapidly evolving threats. By deeply integrating with existing security infrastructure, ITP relentlessly analyzes user behavior and system data to proactively hunt down and neutralize risks before they impact our customers. Join an elite team of full-stack engineers, visionary data scientists, and pioneering ML engineers. We are not just redefining cloud authentication—we are engineering the next-gen security platform that continuously evaluates global threats and constructs a comprehensive, dynamic risk . ITP stands as the uncompromising guardian of trust and security, placing decisive risk prioritization at the core of every authentication and authorization decision. Identity Threat Protection with Okta AI The Staff Software Engineer Opportunity Okta is seeking a Staff Software Engineer to be a driving force on the ITP engineering team. This is your chance to dive into the deep end and solve the most critical and complex security challenges facing enterprises today. We seek a candidate who has not just experience, but a proven record of designing, building, launching, and scaling world-class security products —from concept to global deployment. You will be a technical cornerstone of a team that lives and breathes elegant solutions and

javareactangular
View job →
O
Okta
📍 Toronto• Full-time• From C$160K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is responsible for the foundational data services, systems, and data products for Okta that benefit our users. Today, the Data Platform team solves challenges and enables: Streaming analytics Interactive end-user reporting Data and ML platform for Okta to scale Telemetry of our products and data Our elite team is fast, creative and flexible. We encourage ownership. We expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company. Okta is about to change the cloud computing landscape forever. About the Position This is an opportunity for experienced Software Reliability Engineers to join our fast growing Data Platform organization that is passionate about scaling high volume, low-latency, distributed data-platform services & data products. In this role, you will get to work with engineers throughout the organization to build foundational infrastructure that allows Okta to scale for years to come. As a member of the Data Platform team, you will be responsible for designing, building, and deploying the systems that power our data analytics and ML. Our analytics infrastructure stack sits on top of many modern technologies, including Kinesis, Flink, ElasticSearch, and Snowflake. We are looking for experienced Software Engineers who can help desi

javaawskubernetes
View job →

About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches

REMOTEpythonawsazure
View job →
HI
HP IQ
📍 San Francisco• $149.9K – $270K/yr
11 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Senior Platform Engineer at HP IQ, you will help build and evolve the infrastructure, tooling, and shared platform capabilities that enable our engineering teams to develop and operate reliable, secure, and scalable services across cloud and edge environments . You will work closely with application, services, AI/ML, and security teams to improve developer velocity, production readiness, reliability, and operational efficiency across a heterogeneous infrastructure footprint. What You Might Do Design, build, and maintain shared infrastructure and platform capabilities across cloud and edge environments. Build automation and self-service tooling that improves engineering velocity and operational consistency. Develop and maintain Infrastructure-as-Code, deployment workflows, and environment provisioning. Partner with engineering teams on production readiness, including reliability, security, observability, scalability, and recovery. Improve monitoring, alerting, incident response, and operational tooling across distributed environments. Automate repetitive operational t

pythonkubernetesai
View job →
O
1mo ago

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities Co-design future hardware for programmability and performance with our hardware vendors Assist hardware vendors in developing optimal kernels and add support for it in our compiler Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory h

pythonawsrest
View job →
J
17 days ago

Backend Engineer (Senior Level) - SDE IV We're looking for a Senior Backend Engineer to lead the architecture and evolution of backend services that deploy and serve machine learning models in production. You'll work closely with ML Engineers, Platform, and Product teams to build scalable, reliable systems and drive technical direction across multiple teams. What You’ll Do Design and drive the long-term architecture of backend services for biometrics and ML model serving. Collaborate with core platform and backend teams on organization-wide architectural initiatives. Partner with business and engineering teams to design and deliver cross-cutting platform capabilities. Lead architectural reviews, mentor engineers, and promote engineering best practices. Build and maintain backend services for deploying and serving ML models Monitor service reliability, performance, and scalability in production Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 + Kubernetes Support real-time and batch inference workflows Contribute to CI/CD pipelines and deployment automation What We’re Looking For Strong expertise in backend development using Java and working knowledge of Python. Experience mentoring engineers and driving architectural decisions. Working knowledge of Python, especially for ML-related workflows Hands-on experience with AWS (e.g., DynamoDB, ECS, EC2, Redis, S3, SageMaker) Familiarity with Terraform or other infrastructure-as-code tools, and experience with CI/CD and production monitoring Experience with observability tools (Datadog, New Relic, etc.) Experience with containers and orchestration (Docker, ECS, etc.) Understanding of how ML models are deployed and served in production Experience with Kubernetes Nice to Have Experience with MLOps or ML platform engineering. Experience with asynchronous programming and event-driven systems. Jumio Values: IDEAL: Integrity, Diversity, Empowerment, Accountability, Leading Innovation Equal Opportunities :

REMOTEpythonjavaredis
View job →
S
1mo ago

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data Technology Eva

pythonsqlaws
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur

pythonawsgit
View job →
A
Abbott
📍 United States
12 days ago

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Overview The AI Platform Engineer builds and operates the machine learning and generative AI platform used by teams across Abbott Cancer Diagnostics. You'll own the full model lifecycle in production — data and feature pipelines, training and experimentation, evaluation and promotion, serving, and monitoring — along with the platform services, compute and tooling underneath it. This is hands-on infrastructure work backed by solid platform engineering practice: making inference fast and cheap, making the path from experiment to production repeatable and auditable, and shipping interfaces other engineers can build on — in support of software that ultimately reaches patients. Essential Duties Include, but are not limited to, the following: Build and maintain data, feature, and training pipelines for ML and LLM workloads — ingestion, transformation, fine-tuning, distributed training, and reproducible experiment execution with lineage tracked from dataset and code to resulting model. Implement automated evaluation and promotion gates — performance benchmarks, regression checks, and validation criteria that determine whether a model advances toward production. Automate the model lifecycle end to end through CI/CD and GitOps: packaging, promotion across environments, progressive rollout, and rollback. Build and operate production model-serving infrastructure for LLMs and predictive models, including inference optimization, autoscaling,

pythonjavaaws
View job →
P
Pagerduty
📍 Atlanta• Full-time• From $98K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure

pythonawsazure
View job →
🔔

Get new software engineer ml infrastructure platform salary india jobs by email

Daily job updates · Unsubscribe anytime