Jobiba hiring network

Lead Infrastructure Software Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
Mongodb
📍 Gurugram• Full-time
1mo ago

MongoDB seeks an experienced Senior Software Engineer to help level up IT and play a key role in building a new engineering team. This team works alongside the Internal Engineering department to build enterprise-grade software that enables coworkers to be more effective and efficient. As a Senior Software Engineer, you’ll take ownership of critical internal tools and platforms, drive technical decisions, and contribute at a high level to systems that support essential workflows across the organization. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations: Own the implementation and long-term maintenance of internal tools and platforms Lead technical design discussions and contribute to architectural decisions Write high-quality, maintainable code while setting a strong engineering example for the team Identify and proactively address technical debt, reliability risks, and scalability concerns Collaborate with engineers, product partners, and other stakeholders to prioritize and deliver high-impact work Success Measures: In three months, demonstrate strong ownership of one or more internal systems and contribute significantly to ongoing initiatives In three months, independently deliver meaningful improvements that reduce manual effort or operational friction In six months, have driven technical improvements or new tooling that measurably improves internal engineering efficiency, reliability, or velocity Our ideal candidate: Has 6+ years of professional software engineering experience Deep experience with at least one modern programming language (Python, Go, Rust, etc.) Strong technical judgment and the ability to independently solve complex engineering problems Excellent communication skills and comfort collaborating across teams and disciplines Comfortable working with modern infrastructure and delivery systems, including containerized applications, Kubernetes, and CI/CD tooling (e.g., Drone.io or simil

pythonreactmongodb
View job →
M
Mongodb
📍 United States• Full-time• From $151K/yr
1mo ago

We’re looking for a Senior Engineering Manager who is ready to lead through ambiguity and improve how software gets built at MongoDB. This role leads teams focused on developer productivity, with an emphasis on measurable improvements to the software development lifecycle. This role can be based remotely in the United States. The Team The AXIS team (AI, X-functional tools, Insights, and Signals) sits within Developer Productivity and is responsible for overseeing the metrics and observability infrastructure of our expansive developer environment to help build a strong data-driven culture. You’ll also be a key partner in building the agentic ecosystem for AI-driven development across engineering. Candidate Profile We’re looking for an experienced leader with a passion for solving the big challenge of measuring developer productivity and providing the actionable signals that help teams improve their performance. They should be comfortable working collaboratively with other leaders and partners across our Engineering and Data teams in maximizing the use of data for insights and AI enablement. The right candidate for this role will have 4+ years of experience managing software engineers, including hiring, performance management, growth planning, and compensation; required for external candidates and preferred for internal candidates 8+ years of hands-on software engineering experience building and operating production systems; experience in developer tooling, platform engineering, observability, or data engineering is a strong plus Demonstrated the ability to lead through ambiguity, work across team boundaries, and deliver outcomes without close supervision Strong customer orientation and sound judgment in finding practical, high-leverage solutions Experience working with systems involving analytics, data pipelines, and metrics platforms Experience with AI tools development and enablement efforts Strong technical judgment, including the ability to evaluate t

mongodbawsazure
View job →
M
Mongodb
📍 Alberta• Full-time• From C$191K/yr
1mo ago

We’re looking for a Senior Engineering Manager who is ready to lead through ambiguity and improve how software gets built at MongoDB. This role leads teams focused on developer productivity, with an emphasis on measurable improvements to the software development lifecycle. This role can be based remotely in Canada. The Team The AXIS team (AI, X-functional tools, Insights, and Signals) sits within Developer Productivity and is responsible for overseeing the metrics and observability infrastructure of our expansive developer environment to help build a strong data-driven culture. You’ll also be a key partner in building the agentic ecosystem for AI-driven development across engineering. Candidate Profile We’re looking for an experienced leader with a passion for solving the big challenge of measuring developer productivity and providing the actionable signals that help teams improve their performance. They should be comfortable working collaboratively with other leaders and partners across our Engineering and Data teams in maximizing the use of data for insights and AI enablement. The right candidate for this role will have 4+ years of experience managing software engineers, including hiring, performance management, growth planning, and compensation; required for external candidates and preferred for internal candidates 8+ years of hands-on software engineering experience building and operating production systems; experience in developer tooling, platform engineering, observability, or data engineering is a strong plus Demonstrated the ability to lead through ambiguity, work across team boundaries, and deliver outcomes without close supervision Strong customer orientation and sound judgment in finding practical, high-leverage solutions Experience working with systems involving analytics, data pipelines, and metrics platforms Experience with AI tools development and enablement efforts Strong technical judgment, including the ability to evaluate tradeoffs, i

mongodbawsazure
View job →
G
15 days ago

Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl

pythonlinuxai
View job →
G
15 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Debug Validation Lead will drive post-silicon debug and validation activities for next-generation AI compute silicon and systems. The role is responsible for leading teams focused on identifying, reproducing, analysing and resolving complex silicon, firmware and system-level issues during bring-up, characterization and product readiness. This position combines deep technical debugging expertise with strong cross-functional collaboration across multiple engineering disciplines. The role will work closely with architecture, RTL, firmware, software and systems teams to improve debug methodologies, accelerate issue resolution and strengthen validation coverage. The role will work closely with architecture, RTL, firmware, software, systems and platform teams to improve debug methodologies, accelerate issue resolution and strengthen validation coverage. The Team The Post-Silicon Debug and Validation team sits within the Architecture and Validation organisation and is responsible for bring-up, debug and validation of Graphcore silicon and systems. The

pythongitlinux
View job →
M
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed Engineers on our engineering team who want to work at the intersection of deep infrastructure work and direct customer impact. As an FDE, you'll partner with leading AI companies and foundation labs on cloud architecture, networking, storage, containerization, sandboxing, and more — helping them design and ship production infrastructure on Modal's platform. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the infrastructure stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and deploy massive-scale production workloads on Modal Lead technical discovery and architect

awsazuregcp
View job →
D
1mo ago

As Engineering Manager for Code Security, you'll lead a team of engineers building Infrastructure as Code and Secrets protection - one of the fastest-growing areas inside Datadog's security business, with a strong roadmap and real customer problems to solve. You'll partner closely with Product Management and User Experience to shape how the team delivers, stay hands-on with the technical work, and bring AI deeply into how your engineers build. The organization is still young, which means real room to shape its direction and grow into broader scope as it does. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead and grow a team of engineers building Datadog's Infrastructure as Code and Secrets security products Own the team's roadmap, staffing, and delivery, partnering closely with Product Management and User Experience Stay hands-on: contribute to code review, weigh in on design decisions, and participate in the on-call rotation Bring AI deeply into how the team builds, from adopting agentic coding tools to rethinking workflows around them Coach engineers at every level, giving direct feedback and helping them grow their scope and careers Keep the team's projects and programs organized as priorities shift across a fast-growing area Who You Are: You've managed software engineers for 2+ years, with a track record of shipping through your team You bring a strong technical background you can draw on in day-to-day roadmap and design conversations You're comfortable staying close to the work - reviewing code and unblocking technical decisions alongside your team You've integrated AI agents into how you and your team deliver code, not just experimented with them You keep projects and programs organized, even as priorities and scope shift &n

aigorust
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri

awsrestai
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Direct Surfaces Team’s mission is to make it delightful for users to manage their money. We’re rebuilding and unbundling the core of Stripe so that no matter what the business, the location or use-case there is a single place that can fulfil all financial needs: Stripe Treasury . This is a 0 → 1 new business at Stripe so if that sounds exciting we’d love to hear from you! What you’ll do Full-stack engineers at Stripe are comfortable working on new products under fluid conditions, seamlessly balancing tactical and strategic considerations. Responsibilities Scope and lead technical projects, laying the groundwork for our products to iteratively evolve and scale Design, build, expand and maintain UIs their related APIs and services Work with our partners to design and launch new features and capabilities Align our technical decisions with Stripe’s broad strategic initiatives, while also advocating for needs specific to emerging new businesses Work with engineers across the company to understand when existing infrastructure can be leveraged vs. when building a bespoke solution is prudent Develop and execute against both short- and long-term roadmaps. Make effective tradeoffs that consider business priorities, user experience, and a sustainable technical foundation Who you are We’re looking for fullstack software engineers with experience building scalable products who have an eye for detail and are hap

reactaccounting
View job →
B
15 days ago

At Breeze, we're building the AI-powered infrastructure layer for global commerce, making it radically simpler for businesses to sell, get paid, and operate across markets. We go far beyond traditional payment processing. Breeze combines global payments, AI, stablecoins, and a Merchant of Record-like model to take on the complexity businesses typically manage themselves, including compliance, risk, fraud, chargebacks, reconciliation, and customer support. Our goal is simple: let businesses focus on building and selling great products while Breeze handles the complexity behind getting paid. Backed by Sequoia Capital , Multicoin Capital , and The Chainsmokers , Breeze is a successful, rapidly growing, and exceptionally well-capitalized company. We have the runway to think long term while remaining early enough that every person joining today can have a meaningful impact on what we build. We are hiring a Staff Machine Learning Engineer, Risk! As our Staff Machine Learning Engineer, Risk, you'll lead the evolution of our ML platform for payment risk, building the production-grade capabilities behind feature engineering, model training, deployment, monitoring, and continuous improvement. Risk decisions sit at the center of our business, and you'll own how those models get built, shipped, and kept healthy. This role reports to the CTO. You'll work closely with Risk, Software Engineering, and Data Engineering, and you'll be the senior technical voice for ML on the risk team. We're looking for someone who thrives in fast-moving environments, wants meaningful ownership, and is excited to build rather than simply maintain. What You'll Do Design and build ML infrastructure for payment risk detection, using Databricks as the core platform, in close partnership with software and data engineers. Bring structure to the team's ML environment: feature pipelines, versioning, job orchestration, and monitoring. Design and productionize models rather than just prototype them, including

machine learningaigo
View job →
G
15 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role combines deep technical expertise with people leadership responsibilities, including team development, prioritisation, mentoring and delivery coordination across multiple projects and stakeholders. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug complex issues, optimize workloads and continuously imp

pythonlinuxai
View job →
G
15 days ago

Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl

pythonlinuxai
View job →
G
15 days ago

Senior -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to deb

pythonlinuxai
View job →
G
15 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

pythonaigo
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The UI Platform team, known as Sail, builds the foundations that support Stripe’s Dashboard and many of Stripe’s other user interfaces. The Sail Services team provides backend-for-frontend services that help product teams build consistent, scalable user experiences. The team is defining its mission, roadmap, and technical vision as the UI platform takes on a more central role in how Stripe builds user interfaces. We work with engineers, product managers, and designers across Stripe to connect the design system, software development kits, and backend services into a consistent, surprisingly delightful user experience. What you’ll do As a Staff Software Engineer on the UI Platform team, you’ll help shape the technical direction for the services that support Stripe’s user interfaces. You’ll combine hands-on engineering with technical leadership, working with backend engineers and cross-functional partners to deliver platform capabilities that help Stripe build consistent, reliable user experiences. Responsibilities Define a cross-team technical strategy, architecture, and multi-year roadmap for the backend services that power Stripe’s merchant-facing experiences Lead the design and delivery of scalable backend services that give product teams a consistent foundation for building reliable user experiences across the Dashboard, mobile applications, and other interfaces Build reusable platform capabilities that help product teams develop, operate,

🔔

Get new lead infrastructure software engineer jobs by email

Daily job updates · Unsubscribe anytime