Jobiba hiring network

Lead Infrastructure Software Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

G
Gitlab
📍 India• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. About the role As a Staff Backend Engineer, you will provide technical leadership across your team and adjacent teams, solving the highest-scope and most complex problems in your area. You will lead large, cross-cutting backend initiatives, drive our modular architecture strategy, and define the standards that let teams move faster without compromising quality, security, reliability, or operability. This is a technical leadership role that combines deep backend expertise, systems judgment, product judgment, and influence across organizational boundaries. You will work with Product, Frontend, Infrastructure, Security, Data, Engine

REMOTEsqlpostgresqlkubernetes
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Container runtimes were designed for general-purpose software workloads. AI inference is not a general-purpose workload. Running large models at production scale exposes cracks in every layer of the container stack: runtimes unaware of GPU memory constraints, images that take minutes to pull when a model needs to scale to thousands of replicas, and isolation mechanisms that weren't designed for the multi-tenant serving environments that production AI requires. The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far. Baseten owns the entire pipeline, from the moment a developer pushes a model to the moment a request gets a response. That vertical ownership means we can fix these problems at the root. The Runtime Fabrics team is doing exactly that: purpose-building the container runtime and storage layers for AI inference workloads, led by some of the world's top containerd maintainers. As Engineering Manager of the Runtime Fabrics team, you will lead this work, setting technical direction, growing a world-class team of systems engineers, and ensuring the team's output shapes not just Baseten's infrastructure but the open-source container ecosystem at large. If you've contributed to containerd, runc, or related OCI projects and are ready to lead a team solving some of the hardest problems in infrastructure today, we'd love

linuxmachine learningai
View job →
M
Mixpanel
📍 San Francisco• Full-time• Hybrid• From $320K/yr
1mo ago

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About Mixpanel Engineering Mixpanel Engineering is a small, fast-moving team focused on delivering real value to customers. We build powerful AI-powered product analytics while obsessing over clarity, simplicity, and delight. Product innovation drives our business, and product engineering teams own that responsibility. The Product Engineering org is responsible for building our core product that delivers insights to customers. These teams are defining the next phase of our AI-led product experiences and are uniquely positioned to build the next generation of AI-first products. We build the tools product managers rely on to understand users, protect revenue, and scale companies. As we move toward IPO, our work evolves from showing what happened to explaining why. This shift unlocks deeper insight, better decisions, and the next generation of analytics. About the Role The Director of Engineering, reporting directly to the CTO, will be responsible for all of Product Engineering. You'll own the core analysis experience that product managers and engineers rely on daily, the growth infrastructure that brings new builders into the product and keeps them there, and the platform capabilities that let teams test and ship with confidence. Success is measured by revenue and retention. What You'll Do Lead multiple engineering teams across different product areas. This role is for someone who develops people seriously and holds a high bar for execution. Build an AI first product analytics for the next generation of software Hire well, coach deeply, and build a culture where engineers feel ownership and pride in what t

awsgitrest
View job →
D
Datadog
📍 Massachusetts• Full-time• From $244K/yr
1mo ago

Datadog’s Cloud Networks team designs, builds, and maintains the production network infrastructure that powers everything built on top of our platform across AWS, GCP, Azure, and beyond. In this role, you’ll set technical direction for how we scale our multi-region, multi-cloud network footprint while keeping reliability and performance high. You’ll partner closely with internal teams and Cloud Service Providers to troubleshoot complex connectivity issues, integrate new networking capabilities, and improve the foundations our engineers and customers rely on. This is a high-impact opportunity to drive meaningful improvements in scale, resiliency, and cost efficiency. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design, build, and operate cloud network infrastructure across AWS, GCP, Azure, and Neoclouds in a multi-region environment. Own connectivity between clouds, customers, and developers—ensuring scalable, secure, and reliable network paths. Set clear technical direction for expanding data centers and evolving the network while maintaining stability and performance. Improve cross-site and cross-region connectivity patterns to support Datadog’s growing platform needs. Lead deep investigations into latency, packet loss, and connectivity failures – from pcap and path analysis through to escalations with cloud providers that may originate from customer support Identify and deliver network-related efficiency and cost-saving opportunities that positively impact business health. Who You Are: You have deep networking expertise. You understand BGP, route policies, path selection, prefix advertisement, and what breaks in large-scale networking. You have substantial experience designing, building, and evolving large-scale Software-Defined Networks—inclu

awsazuregcp
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pr

mongodbawsazure
View job →
M
Mongodb
📍 New York City• Full-time• From $157K/yr
1mo ago

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workf

mongodbawsazure
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. As the Site Reliability Engineering Manager for SLS, you will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll help grow and lead a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. We are looking to speak to candidates who are based in Cork for our hybrid working model. Responsibilities Build and lead a team of 6-8 engineers, fostering a positive culture, handling career growth and performance conversations, and proactively removing blockers Define and drive a clear technical vision and comprehensive roadmap for our multi-tenant distributed storage systems, balancing long-term strategic infrastructure goals with immediate engineering needs Contribute through hands-on technical work, such as leading architectural design reviews, reviewing PRs, and stepping in to guide the team through complex operational challenges Act as the primary liaison for the Storage Layer Services SRE team, collaborating closely with other engineering leaders to ensure platform alignment and manage stakeholder expectations You may be a good fit if you Have 10+ years of experience working on software and operating distributed systems, with 2+ years managing engineering teams Possess a customer-focused mindset, treating internal developers as your primary users Value efficiency in processes and operations, and have a track record of optimizing team workflows Pref

mongodbawsazure
View job →
O
Okta
📍 Toronto• Full-time• From C$146K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and resiliency improvements for the Auth0 product at Okta. Here you'll be working with some of the most advanced technology in the space, helping to streamline and secure billions of access requests a year. In this role, you will work closely with architects, platform team members, and product engineers to build the testing infrastructure that keeps Auth0 performant at scale — including the frameworks, tooling, and realistic datasets that make that testing meaningful. The ideal candidate is passionate about software quality and architecture, a self-starter, intellectually curious, and brings deep experience with performance testing, load testing frameworks, dataset generation, performance analysis, monitoring tooling, and chaos engineering. What you’ll be doing Collaborate with architects, tech lead, product owners, security and operations engineers to implement best practices related to performance and resiliency Communicate and organize cross-team projects with high business

javascripttypescriptjava
View job →
L
Lyft
📍 Toronto• Full-time• From C$96K/yr
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft Urban Solutions (LUS) is North America’s leader in micromobility. We own, operate, or provide hardware and software solutions for bikeshare and scootershare programs in 50+ global markets including Montreal, Toronto, London, New York City, Mexico City and others. Our rapidly growing active fleet includes state-of-the-art charging stations, electric bikes, and scooters, and services hundreds of millions of rides per year. Data and analytics are at the heart of Lyft's products and decision-making. You will play a key role in shaping the future of bikeshare by leveraging data to improve the performance of our bikeshare and scooter markets. A successful candidate thrives in a dynamic and collaborative environment, has a natural curiosity, and isn’t afraid to dive deep. In this role, you will collaborate closely with Operations, Policy, Engineering, Product and Finance teams to drive data-informed strategies that align our operations and products with city transportation goals and user needs. You will work in a fast-paced environment where analytical insights directly impact decisions ranging from pricing and product features to long-term investments in bikesharing infrastructure. We’re looking for a passionate and driven Data Analyst to tackle some of the most complex and impactful challenges in micromobility. If you’re excited about shaping the future of urban mobility through data, we’d love to hear from you. Responsibilities: Partner with Product, Engineering, Policy, Operations, Finance and other cross-functional stakeholders on initiatives to conduct deep-dive analyses to root cause issues and propose solutions Develop frameworks, business logic and scalable processes to streamline reporting and drive decision-making Forecast operational requirements and investments needed to

pythonsqlai
View job →

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

awsazuredocker
View job →
G
Gitlab
📍 United Kingdom• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. We're looking for a technical engineering manager to lead a team building and operating large-scale distributed systems. The ideal candidate combines deep technical expertise with strong operational instincts and a passion for building high-performing teams. You will lead the Dedicated Platform Integration team, responsible for the integration of third-party and modular features in the GitLab ecosystem, for Dedicated platform hosted SaaS. Key Responsibilities: Enable effective execution with Quality and Speed, in partnership with the team’s Product Manager Ensure 3+ 9s availability of Dedicated infrastructure, ensuring secu

REMOTEgitrestai
View job →
G
Gitlab
📍 Bengaluru• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role This role is not a typical infrastructure operations role. As Manager, Engineering, Git & Gitaly Operations, you'll lead the team responsible for the infrastructure layer underneath Git at GitLab: the open source Git project itself, Gitaly, GitLab's Git storage and remote procedure call layer, and the services that make Git perform at enterprise scale. Git and Gitaly are core to GitLab. If they go sideways in production, they affect everything downstream, including continuous integration, code review, and deployments. This role starts with genuine subject matter expertise in Git technologies and exten

REMOTEgitrestai
View job →

Datadog (NASDAQ: DDOG) is looking for a Product Strategy and Corporate Development Lead to drive Product vision, M&A, and venture investments for Datadog. You will work directly with Datadog's founders, partnering with Datadog's global engineering and product leaders to identify and execute transactions that expand our platform into new markets. This is not a traditional Corp Dev seat. You will operate on a small, high-autonomy team where technical depth matters as much as deal execution. You can expect to focus on the EMEA landscape, which is one of the most dynamic ecosystems in enterprise software right now, and you will be at the center of it. At the same time, you'll be working with our Product and Engineering leaders based in Europe, who will depend on you to be the fabric between our NYC and Paris HQs. Together, you'll be expected to go deep and autonomously explore new areas of expansion for us. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: (Describe role responsibilities here) Be close to our Product & Engineering leaders to discover new areas of interest and acceleration before an M&A target is even identified, and stay ahead of market trends in AI/ML, observability, security, and cloud infrastructure to inform Datadog's growth strategy Identify and evaluate M&A targets and venture investment opportunities across European markets, with a focus on AI-native and infrastructure companies Build deep relationships with founders, VCs, and accelerators across the Paris, London, and Iberian ecosystems to generate proprietary deal flow Lead end-to-end due diligence: technical product assessments, financial modeling, valuation analysis, and integration planning Partner with senior engineering and product leaders to assess techni

restaigo
View job →

About the Team The Solutions Engineering team consists of trusted technical advisors who help organizations adopt OpenAI’s technology safely, effectively, and responsibly. We partner closely with customers, Sales, Product, Engineering, Research, and Security to translate frontier AI capabilities into practical workflows that create measurable impact. Cybersecurity is one of the most important areas where AI can help. As frontier models become increasingly capable of reasoning across code, logs, infrastructure, vulnerabilities, and security evidence, customers need expert guidance to evaluate and deploy these systems safely. Our Field Security Specialists help security leaders and practitioners apply OpenAI models, APIs, Codex, agentic workflows, and emerging cyber capabilities to real defensive challenges. About the Role We are looking for a Manager of Field Security Specialists to build and lead our customer-facing cyber solutions engineering function. This is a hands-on leadership role for someone who can develop an exceptional team while remaining close to the technology and our customers. You will establish the operating model for the function, raise the quality of specialist engagements, and personally support our most strategic and technically complex customer opportunities. You will work across CISO-level strategy, practitioner-level cybersecurity challenges, and hands-on solution design. Your team will help customers explore workflows including application security, secure software development, vulnerability management, threat modeling, cloud and identity security, SOC operations, detection engineering, incident response, and security validation. This is not an internal CISO or corporate incident-response role. It is a customer-facing leadership opportunity focused on helping defenders achieve safe, measurable outcomes with frontier AI. In this role, you will: Hire, coach, develop, and lead a high-performing team of Field Security Specialists. Establish the

awsci/cdgit
View job →
T
16 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. At Tenstorrent, we build open, state of the art compute for real workloads and real developers. You will own CPU core-level testbench development and verification, shaping how our out-of-order RISC-V CPUs behave in silicon. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You bring 8+ years in CPU verification, CPU testbench development, or closely related digital design. You have deep hands-on experience building and owning CPU core-level testbenches, not just using existing environments. You know high-performance out-of-order CPU microarchitecture in depth. You are comfortable developing testbench infrastructure in CVM methodology, with UVM experience as a strong plus. You work comfortably across RTL, waveforms, logs, regressions, and cross-functional debug with design, DV, emulation, and post-silicon teams. You are comfortable using AI-assisted verification workflows to improve debug, stimulus creation, and coverage analysis, while applying strong engineering judgment to validate results. What We Need Lead hands-on CPU core-level testbench development for hi

pythonawsgit
View job →
🔔

Get new lead infrastructure software engineer jobs by email

Daily job updates · Unsubscribe anytime