Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The SRE Leadership Team The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible—it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders. What You'll Be Doing Lead the SRE team's technical direction , translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems Build infrastructure resilience , designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency Champion reliability best practices , establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts Mentor and develop SRE talent , elevati
Jobiba hiring network
Lead Infrastructure Software Engineer Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead infrastructure software engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
👋 Welcome to Glide! At Glide we’re reimagining the banking experience for the modern world . Our embedded fintech platform empowers legacy financial institutions, like community banks and credit unions, to pioneer novel digital experiences for their customers. You’ll be joining an all-star team with engineering, product, and growth experience from Stripe, Google, and Amazon. We’re looking for a talented early career Fullstack Engineer to help us build our platform. We’re bringing a new perspective to the decades-old financial world , and we’re hoping you can help us do that! Why You’ll Love This Role (Especially if you’re a new grad): A unique opportunity to work side-by-side with Glide’s founders and senior engineers, learning directly from industry veterans. Gain hands-on, end-to-end experience across the full stack: from architecture design to deployment. Be part of a small, fast-moving startup, where your code has immediate impact. Get mentorship and exposure to real-world engineering best practices, product development, and startup culture. Your Responsibilities: Move seamlessly between frontend and backend to build a well-tested, secure frontend for our core web product. Create trustworthy, safe user experiences by building interfaces that are simple, reliable, and performant using tools like Typescript, React, Node.js and NextJS . Design a scalable architecture that can serve hundreds of thousands of end users. Bring designs to life through beautifully-crafted code. Articulate a long-term technical direction and vision for maintaining and scaling our web product suite. Lead frontend and backend infrastructure & tooling for an ambitious product roadmap. Need-to-Haves: Experience with Javascript and tools like Typescript, React, Node.js and NextJS . Experience with modern, responsive HTML & CSS . Experience using data fetching libraries like React Query/Tanstack or tRPC to synchronize client and server data. Excellent understanding of software engineer
👋 Welcome to Glide! At Glide we’re reimagining the banking experience for the modern world . Our embedded fintech platform empowers legacy financial institutions, like community banks and credit unions, to pioneer novel digital experiences for their customers. You’ll be joining an all-star team with engineering, product, and growth experience from Stripe, Google, and Amazon. We’re looking for a talented early career Fullstack Engineer to help us build our platform. We’re bringing a new perspective to the decades-old financial world , and we’re hoping you can help us do that! Why You’ll Love This Role (Especially if you’re a new grad): A unique opportunity to work side-by-side with Glide’s founders and senior engineers, learning directly from industry veterans. Gain hands-on, end-to-end experience across the full stack: from architecture design to deployment. Be part of a small, fast-moving startup, where your code has immediate impact. Get mentorship and exposure to real-world engineering best practices, product development, and startup culture. Your Responsibilities: Move seamlessly between frontend and backend to build a well-tested, secure frontend for our core web product. Create trustworthy, safe user experiences by building interfaces that are simple, reliable, and performant using tools like Typescript, React, Node.js and NextJS . Design a scalable architecture that can serve hundreds of thousands of end users. Bring designs to life through beautifully-crafted code. Articulate a long-term technical direction and vision for maintaining and scaling our web product suite. Lead frontend and backend infrastructure & tooling for an ambitious product roadmap. Need-to-Haves: Experience with Javascript and tools like Typescript, React, Node.js and NextJS . Experience with modern, responsive HTML & CSS . Experience using data fetching libraries like React Query/Tanstack or tRPC to synchronize client and server data. Excellent understanding of software engineer
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta Identity Governance (OIG) organization is looking for a Software Engineer Manager to join our team — OIG is Okta’s Identity Governance and Administration solution that is directly responsible for one of the most critical and visible workflows in enterprise identity: how employees request, approve, and gain access to the resources they need. Opportunity As a Software Engineer Manager on the OIG team, you will manage the team that is responsible for designing and building Identity Governance product features — spanning the across different access governance personas. You will lead a team of elite engineers, grow the team, and drive impact and customer satisfaction. You will drive engineering best practices, productivity enhancement, and make our elite engineering team shine. You will also work collaboratively across engineering, product, and design to deliver features that are secure, scalable, and delightful to use. This is a rare opportu
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We’re looking for a principal software engineer to lead the next generation of data infrastructure at Pinterest which powers mission critical big data and AI applications. You’ll be working on some of the most exciting big data and AI open source technologies (Flink, Spark, Kubernetes, etc.), at the scale of exabytes of data to help Pinners discover and do what they love. What you’ll do: Lead the strategy and technical direction of Pinterest’s data infrastructure for big data and AI applications Build and scale data infra frameworks and infrastructure to process petabytes-scale datasets, including compute engines, job management, resource management, scheduling and remote shuffling Work with internal customers on critical business use cases that rely on big data Provide thought leadership to the entire company on how data should be
Opportunity Overview: We’re looking for a Manager, Platform Engineering that can lead and grow a high-performing engineering team focused on Developer Experience, DevOps, SRE, and Quality. You will own the systems and processes that enable teams to build, test, release, and operate software with high velocity and reliability, driving engineering efficiency and operational excellence across the organization. What you’ll do: Lead a fast-paced, autonomous team of engineers focused on platform engineering, developer experience, DevOps, SRE, and quality engineering Own and drive the internal developer platform strategy and roadmap, improving how engineering teams build, test, deploy, and operate services Create transparency into engineering efficiency and system health through meaningful metrics across delivery, reliability, and quality Enable teams to move faster by improving CI CD pipelines, environments, tooling, and overall developer workflows Provide technical leadership across platform, infrastructure, and reliability, helping teams build scalable and resilient systems Ensure strong engineering practices across release processes, testing, quality, reliability, and security Define and enforce release guardrails, validation standards, and rollback mechanisms to improve production safety Improve environment stability and consistency across development, QA, and pre production environments Drive test strategy and automation maturity to improve overall product quality and confidence in releases Define and implement observability, monitoring, and alerting standards across systems Improve incident detection, response, and RCA practices, ensuring learnings translate into platform and system improvements Drive cloud infrastructure best practices across AWS, containers, and infrastructure as code Foster a culture of ownership, reliability, and continuous improvement within the team Provide innovative solutions for attracting, developing, and retaining top engineering talent I
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Content Platform team at Roblox powers the infrastructure behind every asset used across the Roblox ecosystem—enabling creators and developers to bring their visions to life at global scale. From 3D models and images to videos and audio, our platform manages the complete lifecycle of all assets essential for immersive experiences, supporting one of the largest services in the world at over 100+ million requests per second. Our mission is to deliver a seamless, reliable, and innovative content system that empowers creators, supports record-breaking games, and ensures the highest standards of performance and safety for our community. As the Technical Director for Content Platform, you will lead multidisciplinary engineering teams responsible for the technical and product vision of Roblox’s asset infrastructure. You will own the lifecycle of every asset—from creation and upload to storage, indexing, delivery, and rendering in the game client. Your leadership will be critical in scaling our systems, optimizing distributed infrastructure, and enabling new possibilities for creators and players alike. You Will: Define and drive the long-term strategy, architecture, and priorities for the Cont
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We build the shared infrastructure, tooling, and lab environments that let teams test quickly, understand failures, and ship with confidence. About the Role As a Systems Integration Manager , you will lead the team responsible for device validation infrastructure, test automation, developer tooling, and lab operations. This is a player-coach leadership role: you’ll set technical and operational direction, build and develop a team of engineers and lab operations professionals, and stay close to the architecture and hardest systems problems. You will partner closely with device software, OS, firmware, hardware, reliability, QA, and release infrastructure teams to define validation strategy, improve release readiness, and ensure our test environments and quality signals scale with the product. Because this is a new category of devices, you’ll have the rare opportunity to build the validation foundation early—shaping the systems, standards, and operating model that will support products from prototype through launch. We’re looking for a leader who combines strong technical judgment with people leadership, operational rigor, and experience building reliable systems for complex hardware-software products. This role is b
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps our users extend their online presence to the physical world. The Terminal team’s mission is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments. With Terminal, businesses can unlock in-person payments use cases that are right for their business model—whether it’s creating a superb retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. We’re looking for an experienced product manager to lead and shape the future of the software platform that powers Stripe Terminal’s devices portfolio. Working closely with your engineering partners you will design, build and launch device software capabilities that delight users and differentiate Stripe’s solutions in the market. Working closely with our hardware experts you will build the multi-year strategy for how Stripe will continue to enhance the scalability, reliability, usability, and market differentiation of the software capabilities of our in-person commerce devices. In this role, you will work with a spectrum of users, from our largest platforms to small start ups as well as external technology partners to deeply understand user needs and market trends. You will obsess over stability, scalability, and expanding our product while delivering the best in person experiences in the world. What you'll do: Set a motivating multi-year vision for device softw
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. We are looking for a manager to lead the Documentation (Doc) Tools team supporting the Builder Experience department. Reporting to the Information Development Director, you will be leading a team of engineers who support the help.okta.com and developer.okta.com sites as well as the teams who contribute and create content for those sites (Information Development, Builder Advocacy, and Content Strategy). What you’ll be doing Lead a team of engineers responsible for building and maintaining the help.okta.com and developer.okta.com content sites. You’ll lead the adoption of AI across your team’s day-to-day deliverables, integrating it thoughtfully to improve your internal stakeholder’s productivity, quality, and speed. Work closely with the Doc Tools leads to develop and implement our roadmap. Recruit great engineers to accelerate our work. Mentor, retain, and develop engineers as they advance in their own careers. Establishing clear priorities, expectations, and accountability for individuals and team Collaborate with the Doc Tools leads to assist in delivering projects on the Builder Experience roadmap. What you’ll bring to the role Have managed software engineering teams with end-to-end ownership over their technical stack and production performance. Hands-on experience with AI-assisted development tools such as GitHub Copilot, Codex and Claude, with the ability to integrate them effectively into day-to-day engineering workflows. Are excited to b
We are now looking for a dynamic business leader to grow NVIDIA's Host Networking business for AI infrastructure with AI Labs and Hyperscalers! This leader will drive strategic direction, customer engagement, and multi-year growth for networking products such as NVIDIA DPUs, SuperNICs, and their associated software and ecosystem. Success in this role will be measured by the level of adoption and integration of our Host Networking products with our end customers' workflows and workloads. Success is contingent upon building trust with executives, architects, product leaders, and platform teams across NVIDIA and our largest customers. This leader will lead the go-to-market motion, connecting customer AI factory needs to NVIDIA's networking portfolio and aligning product, sales, engineering, architecture, marketing, and partner teams to secure design wins and scale deployments. What you'll be doing: Identify, develop and close strategic design wins for DPU and SuperNIC with top AI labs and Cloud Service Providers! Build and implement the segment sales growth strategy for host networking across hyperscaler and frontier model AI labs building large scale AI infrastructure. Define customer-specific DPU and SuperNIC value propositions and deployment motions, and lead a matrixed team across product, architects, engineering, sales and marketing teams. Promote NVIDIA host networking products externally and internally, positioning their value for AI workloads and other infrastructure products from NVIDIA, in a collection of use-cases in Networking, Security and Storage. Build a robust opportunity pipeline with segment sales and account teams, including account mapping, customer requirements, proof points, executive engagement, and partner alignment. Track and drive quarterly business reporting, forecast accuracy, design-win progress, roadmap asks, and
AI/ML – Investment Services A Career with Point72's AI/ML – Investment Services Team The AI/ML – Investment Services team at Point72 spearheads the development of cutting-edge AI solutions that seek to transform our business processes and enhance enterprise intelligence. The team aims to bridge the gap between business challenges and technological innovation, collaborating with stakeholders across the firm and leveraging expertise in generative AI, data engineering, and machine learning. WHAT YOU'LL DO Build and scale core backend services and platforms that power generative AI applications and data infrastructure used across the firm’s investment workflows Design and implement high-throughput, low-latency data pipelines to ingest, normalize, and serve both structured and unstructured data Develop robust APIs and microservices to support model inference, feature serving, and downstream applications Integrate generative AI tools and model-serving workflows into production, including embedding stores, retrieval components, and fine-tuning pipelines Optimize system performance, cost, and reliability through profiling, capacity planning, and architectural improvements Implement automated testing, continuous delivery pipelines, monitoring, and incident response practices to maintain production health Partner with data scientists, AI engineers, product owners, and operations to translate models and prototypes into scalable, production-grade solutions Mentor engineers, lead code reviews, and establish engineering best practices for maintainability, security, and observability Own end-to-end delivery, operational runbooks, and metrics-driven measurement of feature impact and system reliability WHAT'S REQUIRED Bachelor’s degree in computer science, software engineering, or a related technical field Minimum 5+ years of professional experience building backend systems and production services Demonstrated experience designing and operating large-scale data engineering pipelines
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Conversion Visibility Modeling team enables a performant ads marketplace and helps prove value to advertisers by connecting Pinterest onsite activity with conversions that happen offsite (both digital and physical) in a privacy-preserving way. As a Machine Learning Engineering Manager on this team, you will lead a hybrid team of ML engineers and backend software engineers to build end-to-end identity and conversion visibility solutions across modeling, serving, and data infrastructure, so advertisers retain accurate, privacy-aware performance visibility as signals fragment and degrade. You will set the technical direction for high-impact ML systems that feed ranking, bidding, measurement, and reporting across Pinterest’s ads stack. What you’ll do: Attract, hire, develop, and lead a hybrid team of ML engineers and backend softwar
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security that could scale to an extreme level of severity. Our work involves: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role As AI agents become more capable at software engineering, and automate more of our internal work, they could become a dangerous cyber threat. People in this role will help OpenAI prepare for security threats from advanced AI agent insiders. In this role, you will: Identify paths by which capable future internal AI agents could compromise OpenAI. Design security controls - focusing on measures with long lead times that benefit from advanced preparation. Stress-test defenses with AI agent evaluations and penetration tests You might thrive in this role if you: Are deeply technical across security and modern infrastructure, and are comfortable digging into the details of operating systems, cloud, containers, CI/CD, or distributed systems. Have strong software engineering skills and enjoy building prototypes yourself. Are interested in engaging with stakeholders and can do so effectively. Bonus: have experience securing cloud infrastructure, and are deeply familiar with core components of the AI stack. Compensation Range: $293K - $405K USD About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe
Get new lead infrastructure software engineer jobs by email
Daily job updates · Unsubscribe anytime