Jobiba hiring network

Senior Infrastructure Engineer Jobs

7,101 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

G
Godaddy
📍 Bulgaria• Full-time
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. As a Senior Site Reliability Engineer, you'll take direct ownership of production services — from initial design through day-to-day operation — while partnering with product, engineering, and security teams to build and maintain business-critical systems. In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs. You will also gain hands-on experience with intelligent tooling in real-world workflows. What you'll get to do... Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana Lead incident response end-to-end — performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation — validating all outputs before use Your experien

pythondockerkubernetes
View job →

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role We're seeking a Senior/Staff Engineer to build and maintain the automation infrastructure that powers the development cycles of our North platform. This engineer will design and implement robust automation systems that enable engineers to efficiently test and validate changes across diverse environments and configurations. This role sits at the intersection of infrastructure and standards. You'll build the systems, frameworks, and culture that allow the rest of engineering to own quality themselves; improving and extending our testing platform by creating the infrastructure that allows engineers to write and execute tests, and enable every engineering team to ship with more confidence. Key Responsibilities Design and implement automation pipelines that support comprehensive testing across multiple environments with varying feature flags and realistic customer data profiles Create intelligent testing agents that simulate real user behavior to validate different configuration combinations Develop and maintain GitHub workflows and actions to automate testing, deployment, and validation processes Manage and optimize H

typescriptpythonaws
View job →

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Infrastructure builds the systems to ship stable, scalable, and efficient services. We're hiring a Senior Technical Program Manager to lead our AI transformation in developer productivity: rethinking how engineers at Lyft work with AI, from tooling and agentic workflows to the platforms that make them possible. This role blends program delivery with product sense: you'll own the roadmap for your area, set priorities, and act as the voice of the customer. Responsibilities: Lead Lyft's AI transformation in developer productivity: shape the strategy for how AI enhances engineering workflows, from adoption of tools like Cursor, Copilot, and Claude Code to building AI-native workflows and agentic tooling for engineers Own the roadmap: sequence the work, make the prioritization calls, and define what success looks like for AI-driven engineering productivity at Lyft Partner with engineering teams to build AI-native workflows that solve real developer pain, not just deploy tools for the sake of it Run programs end to end, from kickoff through delivery, coordinating across infrastructure, developer platforms, and engineering leadership Be the voice of the customer: partner with engineering teams across Lyft, surface their pain points, and feed that back into priorities and roadmaps Define success metrics and adoption goals, gather and document customer requirements, and make sure what ships actually solves the problem Partner with engineering and infrastructure leads to build plans, call out risks early, and keep stakeholders aligned Own program health: track milestones, surface blockers before they slip, and keep decision-makers in the loop Use your technical background to ask sharp questions and build plans the team believes in Share in the team's release oncall rotation Experience: 5+ years in Technic

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Infrastructure builds the systems to ship stable, scalable, and efficient services. We're hiring a Senior Technical Program Manager to lead our AI transformation in developer productivity: rethinking how engineers at Lyft work with AI, from tooling and agentic workflows to the platforms that make them possible. This role blends program delivery with product sense: you'll own the roadmap for your area, set priorities, and act as the voice of the customer. Responsibilities: Lead Lyft's AI transformation in developer productivity: shape the strategy for how AI enhances engineering workflows, from adoption of tools like Cursor, Copilot, and Claude Code to building AI-native workflows and agentic tooling for engineers Own the roadmap: sequence the work, make the prioritization calls, and define what success looks like for AI-driven engineering productivity at Lyft Partner with engineering teams to build AI-native workflows that solve real developer pain, not just deploy tools for the sake of it Run programs end to end, from kickoff through delivery, coordinating across infrastructure, developer platforms, and engineering leadership Be the voice of the customer: partner with engineering teams across Lyft, surface their pain points, and feed that back into priorities and roadmaps Define success metrics and adoption goals, gather and document customer requirements, and make sure what ships actually solves the problem Partner with engineering and infrastructure leads to build plans, call out risks early, and keep stakeholders aligned Own program health: track milestones, surface blockers before they slip, and keep decision-makers in the loop Use your technical background to ask sharp questions and build plans the team believes in Share in the team's release oncall rotation Experience: 5+ years in Technic

V
Vanta
📍 United States• Full-time• Remote
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. You will build the layer that makes Vanta's data actually useful: a single interpretive layer that reads across every source of truth and makes them queryable, reasoned-over, and genuinely illuminating — for EPD, for GTM, and for anyone in the org trying to understand what's happening and why. The EPD Systems team is building the infrastructure Vanta's engineering, product, and design organization depends on to understand itself. We're constructing three interlocking layers: sources of truth at the foundation, an operating system that makes delivery legible, and the intelligence layer that reasons across all of it. This role owns the intelligence layer — the one that doesn't exist yet. This is a builder role. You will ship working things yourself — prototypes, internal tools, agent workflows. You will not hand specs to someone else and wait. Communication isn't a separate deliverable; it's how you learn what to build. What you’ll do as a Senior Product Builder at Vanta: Build the intelligence layer: a cross-source interpretive layer that reads across Vanta's sources of truth and makes them queryable and reasoned-over by EPD leadership, GTM, and beyond Define what to build: scope the problem yourself, make explicit tradeoffs about what to defer, and own the sequence of what gets built and when Ship working things yourself: prototypes, internal tools, agent workflows — using AI as part of how you work, not what you report on Understand the organization: go to the teams whose decisions this layer will serve — GTM, G&A, EPD — and come back knowing what they can't answer today Catch and address AI-specific quality problems — rel

REMOTEai
View job →
V
Vanta
📍 United States• Full-time• Remote
1mo ago

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. You will build the operating system for EPD — the systems, agents, and practices that make engineering, product, and design unreasonably effective at delivery, with Jira as the substrate. The EPD Systems team is building the infrastructure Vanta's engineering, product, and design organization depends on to understand itself and operate well. We're constructing three interlocking layers: sources of truth at the foundation, a shared interpretive layer that reasons across all of them, and the operating system that makes the work itself legible. This role owns the operating system. This is not a traditional PMO or status-reporting role. You build — agents, automation, workflow design, and the practices that make teams want to use the system rather than route around it. What you’ll do as a Senior Systems Designer at Vanta: Design and own workflow and hierarchy across engineering, product, and design in Jira Make sure Jira reflects real practice, and real practice reflects what needs to show up in Jira — in both directions Build the agents and automation that run the system yourself, with AI as part of how you work Drive adoption: go to the teams whose practice doesn't match the substrate today and change that — through conversation, well-built artifacts, or clear instruction that works without you in the room Understand delivery breakdowns at the mechanism level and build fixes that address the root cause, not the symptom Build alongside teammates who are growing into more technical work — raise their ceiling, not just your own output How to be successful in this role: You build and ship, recently, with AI as part of how you work. "

REMOTEai
View job →
E
17 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a hands-on senior leader for our India SecOps team, you will shape and safeguard Everpure’s security posture at the intersection of detection engineering, threat hunting, attack surface management, and incident response. Positioned as a strategic cornerstone in Bangalore, you will empower an elite engineering team, optimize critical SecOps pipelines, and partner cross-functionally across global engineering and infrastructure groups. By driving execution excellence and high team morale, you ensure our enterprise platform and global telemetry remain resilient against evolving threats. WHAT YOU'LL DO Scale & Lead SecOps Operations: Architect, mentor, and grow the India SecOps team to foster an environment of high morale, technical excellence, and rapid execution across detection engineering and incident response. Proactively Manage & Remediate Attack Surface: Own end-to-end Attack Surface Management (ASM) across cloud environments, SaaS applications, endpoints, and secrets management to measurably minimize enterprise exposure and mitigate risk. Optimize Telemetry & Incident Response: Mature SIEM and SOAR automation pipelines to drastically reduce mean time to detect, contain, and respond (MTTD/MTTC/MTTR) while continuously elevating alert fidelity and signal confidence. Drive Cross-Functional Alignment & RCA Postmortems: Lead continuous validation through purple-teaming and incident postmortems alo

awsazureci/cd
View job →
B
Blocktech
📍 Amsterdam• Full-time
1mo ago

About BlockTech BlockTech is an algorithmic trading firm operating at the frontier of global crypto derivatives and spot markets. We trade 24/7 across some of the fastest-moving, most data-rich venues in finance. Crypto remains one of the few markets where a researcher can still meaningfully move the edge: abundant data, novel microstructure, and the shortest possible loop between a research idea and live PnL. We're looking for an experienced Quantitative Researcher to take ownership of that edge and push it further. The role This is a senior, hands-on research seat on our trading floor. You'll own a research agenda end-to-end from hypothesis, dataset construction, feature engineering, model training, backtesting, live deployment, monitoring, and iteration. You'll be trusted to set its direction. You'll work shoulder-to-shoulder with fellow researchers, traders and analysts, shape how we price and trade, and help raise the bar for research across the floor, including mentoring less experienced researchers and influencing the tools and standards the team relies on. You will: Own price-prediction, signal, execution, and anomaly-detection models across crypto derivatives and spot markets from idea to live PnL using state-of-the-art ML Shape our research, backtesting, and trading infrastructure together with engineers, so good ideas reach production quickly and safely Own models in production: monitor live performance, diagnose decay, and iterate on what you ship Set research direction alongside traders, deciding which trades are worth making and why Raise the research bar by mentoring colleagues, reviewing work, and setting standards for rigour What we're looking for 4+ years of hands-on quantitative research and/or applied ML experience, with a track record of models you've taken into production trading live A strong academic foundation in a quantitative discipline (mathematics, physics, statistics, computer science, ML/AI, or similar) Fluency in Python and the modern

pythonaigo
View job →
B
Blocktech
📍 Singapore, Singapore• Full-time
1mo ago

About BlockTech BlockTech is an algorithmic trading firm operating at the frontier of global crypto derivatives and spot markets. We trade 24/7 across some of the fastest-moving, most data-rich venues in finance. Crypto remains one of the few markets where a researcher can still meaningfully move the edge: abundant data, novel microstructure, and the shortest possible loop between a research idea and live PnL. We're looking for an experienced Quantitative Researcher to take ownership of that edge and push it further. The role This is a senior, hands-on research seat on our trading floor. You'll own a research agenda end-to-end from hypothesis, dataset construction, feature engineering, model training, backtesting, live deployment, monitoring, and iteration. You'll be trusted to set its direction. You'll work shoulder-to-shoulder with fellow researchers, traders and analysts, shape how we price and trade, and help raise the bar for research across the floor, including mentoring less experienced researchers and influencing the tools and standards the team relies on. What you'll do Own price-prediction, signal, execution, and anomaly-detection models across crypto derivatives and spot markets from idea to live PnL using state-of-the-art ML Shape our research, backtesting, and trading infrastructure together with engineers, so good ideas reach production quickly and safely Own models in production: monitor live performance, diagnose decay, and iterate on what you ship Set research direction alongside traders, deciding which trades are worth making and why Raise the research bar by mentoring colleagues, reviewing work, and setting standards for rigour What we're looking for 4+ years of hands-on quantitative research and/or applied ML experience, with a track record of models you've taken into production trading live A strong academic foundation in a quantitative discipline (mathematics, physics, statistics, computer science, ML/AI, or similar) Fluency in Python and the m

pythonaigo
View job →
G
Godaddy
📍 United States• Full-time• From $128K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor

pythonkuberneteslinux
View job →
N
Nuro
📍 Mountain View• Full-time• From $183K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro leverages many different bench-top systems to evaluate and regression test different aspects of the software and hardware integration layer. This performance simulation platform includes systems At Nuro, every autonomy code change, from ML model updates to radius of map around the robot to number of evaluated trajectories, must be validated for real-time performance on actual robot compute hardware before it reaches the road. You will own the infrastructure that makes this possible. Our Performance Simulation Platform is a hybrid benchmarking system: physical bench-top rigs running production robot compute (NVIDIA Thor platform), orchestrated by cloud-native infrastructure (Kubernetes, GCP), automated data pipelines feeding performance metrics into BigQuery and Grafana, pre/post simulation magic, custom tracing and profiling tools, and much much more. Engineers across the company rely on this platform daily to an

pythonsqlgcp
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro is seeking a Software Engineer with expertise in large-scale infrastructure, workload orchestration, and data processing to join our ML Infrastructure team . In this role, you will focus on building and evolving the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management to accelerate the Nuro Driver™ development lifecycle. About the Work You will build the foundation that powers Nuro’s model development from experimentation to production. Key responsibilities include: Resource Provisioning & IaC: Scaling automated infrastructure-as-code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments. Intelligent Scheduling: Designing and optimizing workload orchestration to maximize

redisawsazure
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving, and the ML Infrastructure team builds and operates the infrastructure that makes that possible. We own the systems that train the models at the core of the Nuro Driver™ - from distributed GPU training and closed-loop reinforcement learning, to the workflows, orchestration, observability, and cost management that keep the fleet running efficiently. Our work sits directly on the critical path of autonomy development. When a training run stalls, when a pipeline silently regresses, or when GPU utilization slips, it shows up in how fast the rest of the company can ship. We care as much about reliability and operational maturity as we do about raw scale. About the Work Contribute to Nuro’s training infrastructure, spanning multi-generation accelerators, and multi-cluster scheduling and orchestration. Design and operate large-scale data pipelines - batch and strea

pythongcpkubernetes
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro takes a machine-learning-first approach to autonomous driving technology. In an ML-first system, the overall system performance depends heavily on the quantity and diversity of its training and evaluation data. The team plays a crucial role in the advancement of autonomous driving systems by creating a scalable and reliable data infrastructure. This infrastructure is designed to produce training and evaluation data derived from both on-road collected logs and simulation logs. Additionally, the team collaborates closely with system engineers to thoroughly validate the autonomous driving system before its deployment. About the Work Design and develop unified, introspectable, large-scale batch and streaming data pipelines that can ingest and process data across a wide range of use cases relevant to evaluation. Create and implement a storage system capable of accommodating both the large volume and diverse range of e

pythonsqlpostgresql
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role Our software team is growing, and we are looking for talented engineers to join us and be instrumental to one of the following areas: Onboard Systems, Performance, and Devices Platform. Onboard Systems: Our onboard system team’s software engineers provide a reliable and high-performance platform that allows our autonomy teams to integrate their autonomy software and algorithms that work across various self-driving platforms. This work requires close collaboration with our software teams, hardware teams, and systems/safety team to make sure new software and hardware work together safely and reliably, and resolve onboard error and performance problems. Performance: Our Performance team optimizes the performance of Nuro’s AV software, ensuring our vehicles can react quickly and safely to the world around them. The team builds systems and tools for continuous performance analysis, and drives latency reduction and resource efficiency

pythonreactai
View job →
🔔

Get new senior infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime