Jobiba hiring network

Senior Infrastructure Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

P
Prophecy
📍 Bengaluru• Full-time
18 days ago

About Prophecy The leader in AI-native data preparation and analysis, Prophecy is revolutionizing how the world’s top enterprises turn data chaos into reliable insights. We introduce the AI-native data lifecycle (generate, refine, deploy) where our industry leading AI agents and humans work hand-in-hand in visual and document interfaces to analyze, transform and prepare data, to ship trusted insights at enterprise scale. Don’t miss the rocket ship—join Prophecy and build the next data revolution. Position Summary This is a high-impact opportunity to be a senior DevOps engineer in a fast-growing startup, based in Prophecy’s India engineering center. You will own and evolve the foundations that keep our engineering org fast, secure, and cost-efficient at scale — spanning cloud cost management, security DevOps, and CI/CD and engineering operations. You will work with a team of dynamic engineers who take pride in solving complex problems, and you will have the autonomy to set direction in your areas of ownership. The Impact You Will Have Cloud cost (FinOps) Own and evolve our cloud cost optimization program across AWS, Azure, GCP, Databricks, Snowflake and Bigquery building on the programmatic monitoring and controls Analyze billing, asset-inventory, and utilization data across departments to identify wasteful spend and provide actionable insights to optimize it. Develop and maintain cost optimization strategies, roadmaps, and forecasting models Partner with engineering teams to design cost-efficient architectures without compromising scalability or reliability. Build and maintain automation for infrastructure provisioning, scaling, and cost control. Security DevOps Partner with engineering and security to drive our security-hardening program across workstreams such as identity & access governance, secrets & credential lifecycle, cloud access, and CI/CD hardening. Implement and automate guardrails: secrets management, least-privilege access, cr

pythonawsazure
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

REMOTEkubernetesaigo
View job →
P
24 days ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect

REMOTEpythonawsmachine learning
View job →
G
Godaddy
📍 India• Full-time
26 days ago

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability and l

javanodejssql
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng

pythonsqlpostgresql
View job →
O
27 days ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Role Overview As the Okta Research & Design (ORD) Jira Owner & Administrator , nested within the Technical Program Management (TPM) organization, you will define and execute the strategy for how ORD plans, tracks, and reports on engineering work within the full toolchain of Atlassian Jira, Jira Advanced Roadmaps (Plans), Confluence, Rovo, and AI integrations. In this role, you act as a critical bridge between product & technical leadership, agile teams, and cross-functional business units — translating methodology, best practices, and process frameworks into practical workflow improvements that accelerate delivery and reduce manual overhead for a 1,000+ person engineering, product, and design organization. By optimizing tooling configurations, automating reporting, and eliminating fragmented management processes, you will directly accelerate ORD’s developer velocity and strengthen operational efficiency. Core Responsibilities 1. Platform Ownership & Strategic Tooling Direction Strategic Vision: Define, implement, govern an AI-led strategy for utilizing Jira for product planning, tracking, and reporting tooling that supports end-to-end portfolio tracking from high-level corporate objectives down to sprint-level stories. E2E Ecosystem Management: Complete ORD ownership of the toolchain (Jira, Jira advanced Roadmaps, AI integrations, Confluence), overseeing architecture, schema updates, workflow management, screen/notification schemes, cu

sqlawsci/cd
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. The Cloud Efficiency team builds a unified, self-serve cloud efficiency platform along with AI skills and agents that makes spend observable, attributable, governable while driving recommendations and optimization of our cloud spend. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Design, develop, and maintain scalable platform for resource ownership registry, usage attribution, utilization measurement, and cost modeling. Build AI agents, tools and automation to enhance system monitoring, alerting, and root cause analysis. Improve and optimize data ingestion, storage, and query efficiency for cloud utilization, cost and efficiency data at scale. Collaborate with teams across Snowflake to understand attribution and observability needs and implement solutions that improve operational visibility. Contribute to open-source and industry best practices in monitoring and distributed systems monitoring. Ensure high availability, reliability, and performance of team-managed platforms by participating in on-call rotations and incident management. Partner with Finance, Product and Engineering

pythonjavaaws
View job →
G
1mo ago

Location Details: Remote, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... We are looking for an experienced engineer to join the Infrastructure Management and Automation team. In this role, you’ll help shape and build the tools, platforms, and automation that power and manage GoDaddy’s growing infrastructure footprint while leading the design of scalable solutions to complex operational challenges. Our infrastructure platform empowers GoDaddy engineers to help small business customers start, grow, and run their ventures. GoDaddy teams rely on our infrastructure solutions every day, whether for development workflows or hosting production workloads at global scale. If you enjoy solving complex technical problems, designing scalable systems, influencing technical direction, and improving how infrastructure is managed and automated, we’d love to hear from you! What you'll get to do... Lead the design and evolution of GoDaddy’s next generation of infrastructure management platforms and automation systems, taking ownership of complex technical areas from design through delivery and operation Leverage modern engineering practices and AI-assisted development workflows to rapidly design, build, and deliver scalable internal tooling and automation Design and build reliable services and integrations across infrastructure platforms and operational systems, including provisioning, networking, IPAM, DNS, CMDB, and asset management workflows Partner with engineering and operations teams to identify opportunities, shape technical solutions, and improve infrastructure reliability, scalability, utilization, and opera

javascriptpythonjava
View job →

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and

pythonawsazure
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi

pythonjavaaws
View job →
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

pythonjavakubernetes
View job →
V
Vanta
📍 United States• Full-time• Remote
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers independently drive complex technical work, shape the systems and technical decisions within their teams, and enable other engineers to deliver high-quality, scalable solutions. Vanta's product monitors the security posture of thousands of companies, pulling tens of millions of API calls of data per day, pushing information from hundreds of thousands of laptop agents, and running tests against that data continuously to identify potential security threats. Our infrastructure and tooling need to stay ahead of exponential growth in our customer base. As a Senior Software Engineer at Vanta, you'll drive complex projects across our technical stack, contribute to the technical direction of your team, and mentor other engineers. Your past experience will be leveraged to enable and accelerate Vanta's growth. Visit our Vanta Engineering Blog to learn more about what our team is working on! Tests are at the heart of how Vanta continuously monitors security and compliance for our customers. The Test Core team builds the runtime platform that powers these checks. We own how Tests are scheduled and executed, how their results are persisted and exposed, and the systems that keep this runtime reliable as Vanta grows. In this role, you'll work on some of the core systems behind Vanta's Tests platform. You'll tackle problems around the reliability, correctness, and performance of Test execution, evolve the systems and abstractions that allow the platform to scale, and make it easier for other engineering teams to build on the Tests runtime. Many of these problems span multiple systems and teams and require a deep u

REMOTEtypescriptmongodbgraphql
View job →
O
Okta
📍 Bengaluru• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Job Overview: We are looking for a Senior Engineer to join the FGA DevEx team and help evolve our end to end developer experience across both OSS and SaaS. This team owns the SDKs in Go, JavaScript, .NET, Python, Java and other languages, along with CLI workflows, IDE integrations, GitHub automation, developer documentation, and release strategy. All development is done in the open as open source, and we actively welcome and review community contributions. Our guiding principle is One developer experience, many deployment models. As a Senior Engineer, you will take ownership of significant portions of the SDK and tooling ecosystem, ensure high quality implementations across languages, and contribute to a consistent and reliable developer experience. Responsibilities: Maintain and enhance existing SDKs for FGA DevEx in Go, JavaScript, .NET, Python, and Java, leveraging our SDK generator framework. Customize and refine SDK templates and wrappers to ensure consistency across languages and support configuration overrides such as store ID, authorization model ID, headers, and parallelization limits. Implement and improve core SDK features including client credentials authentication flows, robust error mapping, retry logic with jitter, and rate limiting safeguards. Implement advanced capabilities such as BatchCheck, ListRelations, and non transactional write operations with appropriate parallelization and performance considerations. Contribute to the SDK generato

javascripttypescriptpython
View job →
S
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Senior Software Engineer to own key components of our AI native External Observability Platform . In this role, you will contribute to the technical road map for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex technical projects, and skilled at collaborating with the brightest technical minds in the industry. Key Responsibilities Develop and Scale Distributed Infrastructure: Design and implement key components of Snowf

javavueaws
View job →

About BlockTech BlockTech is a fast-paced algorithmic trading firm facilitating global cryptocurrency derivatives and spot trading while expanding into new markets. As we continue to grow rapidly, we are looking for a Software Engineer to join our Foundation team amid our exciting scale-up phase! You will Build & optimize: Design, develop, and maintain high-reliability, low-latency, and high-throughput foundational systems that enable our trading and technology teams to scale efficiently. Ingest & aggregate: Collect trading business data with minimal latency impact and ingest both public and private exchange information into our trading system. Store & stream: Develop and maintain infrastructure for real-time data aggregation and long-term storage, as well as our Kafka-based messaging systems. Collaborate & support: Work closely with multiple teams, assisting them in integrating with and making the most of our foundational systems. Innovate: Drive projects from concept to deployment with full ownership, and explore new tools, frameworks, and approaches to keep our infrastructure best-in-class. The Foundation team develops core software infrastructure (libraries, frameworks, and systems) for BlockTech, solving common problems and lending its expertise to enable other teams to stay focused on their respective domains. They own, develop, and configure a wide variety of critical, high-reliability software, ranging from low-level ultra-low-latency shared memory IPC libraries to high-throughput data buses and data aggregation systems including Kafka, NATS, PostgreSQL, and Iceberg. They work primarily in Rust, but also use Python and SQL. If you thrive on low-level problem-solving, building robust frameworks from scratch, and enabling others to move faster, this role is for you. What We're Looking For Essential: 5+ years of experience as a Software Engineer, with a strong focus on systems-level optimisation and awareness of hardware constraints Proficiency

pythonsqlpostgresql
View job →
🔔

Get new senior infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime