Jobiba hiring network

Lead Software Engineer Infrastructure Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead software engineer infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're looking for Forward Deployed Engineers on our engineering team who want to work at the intersection of deep infrastructure work and direct customer impact. As an FDE, you'll partner with leading AI companies and foundation labs on cloud architecture, networking, storage, containerization, sandboxing, and more — helping them design and ship production infrastructure on Modal's platform. The FDE team today includes world-class software engineers, computational scientists, ML engineers, and former founders. We're looking for people with strong engineering fundamentals, deep curiosity across the infrastructure stack, and energy for working directly with customers on hard problems. You will: Work hands-on with companies like Suno, Lovable, Cognition, and Meta to architect and deploy massive-scale production workloads on Modal Lead technical discovery and architect

awsazuregcp
View job →
R
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance. You will: Design and Implement Observability Solutions : Develop comprehensive monitoring and alerting systems using modern observability tools. Create dashboards and metrics that provide real-time visibility into system health and performance. Implement logging strategies that enable quick problem identification and resolution. Drive Automation and Infrastructure as Code : Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines that enable reliable and consistent deployments. Create self-healing systems that can automatically respond to common failure scenarios. Establish SLOs and SLIs : Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring we maintain high reliability standards while balancing innovation speed. Incident Management and Response : Lead incident response efforts, conducting thorough post-morte

pythongcpkubernetes
View job →
S
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world's largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Revenue and Financial Automation Sub-org within Revenue and Financial Automation: Billing. The Revenue and Financial Automation team at Stripe builds software tools that accelerate the economic and technological growth of global businesses by helping them operationalize their commercial relationships with customers. Our offerings include a billing platform, SaaS analytics, data services, and finance automation products that our customers creatively combine to support various revenue models. Team Matching: exact team matching for one of the subteams will begin during final stages. Please note we may also consider you for different orgs based on your experience, location, etc. More information on our team matching process can be found here. What you'll do We're looking for engineers who want to build the distributed systems, APIs, and backend services that power Stripe's revenue and billing platform. You'll focus on API design and performance, service reliability, and distributed systems challenges, while collaborating across the stack to ship complete solutions for millions of businesses. Responsibilities Scope, architect, and lead technical projects to build and scale distributed backend systems and APIs Design, build, and maintain reliable, high-performance APIs and backend services Own service reliability, including setting performance targets and driving improvements across the stack Debug production issues across distributed services

S
Stripe
📍 Taipei• Full-time
1mo ago

About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps Stripe users extend their online presence into the physical world. The Terminal team’s mission is to make it as easy for businesses to accept in-person payments as the Stripe API has done for online payments. With Terminal, businesses can unlock in-person payments use cases that are right for their business model—whether it’s creating a flagship retail experience, extending their website to a pop-up store, or enabling a mobile point-of-sale at their next event. What we are looking for: As an Android BSP Engineer, you will be responsible for the kernel and driver level system development, which includes building, troubleshooting, and writing automated tests for the Android system on our embedded payments platforms. This team works closely with partner teams throughout the hardware and software product lifecycle, from hardware manufacturing to Android app teams. We also work with external vendors on part selection and initial hardware bring-up. What you’ll do: Bring up new devices and lead debugging and performance tuning exercises that span multiple hardware/firmware/software teams. Design, implement, and maintain drivers and Android services that operate efficiently in a constrained environment and meet the reliability and security requirements of the industry. Own the definition of one or more work streams focused on hardware bring-up, peripheral drivers and communication, and power and performance management and opti

javaaikotlin
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri

awsrestai
View job →

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Direct Surfaces Team’s mission is to make it delightful for users to manage their money. We’re rebuilding and unbundling the core of Stripe so that no matter what the business, the location or use-case there is a single place that can fulfil all financial needs: Stripe Treasury . This is a 0 → 1 new business at Stripe so if that sounds exciting we’d love to hear from you! What you’ll do Full-stack engineers at Stripe are comfortable working on new products under fluid conditions, seamlessly balancing tactical and strategic considerations. Responsibilities Scope and lead technical projects, laying the groundwork for our products to iteratively evolve and scale Design, build, expand and maintain UIs their related APIs and services Work with our partners to design and launch new features and capabilities Align our technical decisions with Stripe’s broad strategic initiatives, while also advocating for needs specific to emerging new businesses Work with engineers across the company to understand when existing infrastructure can be leveraged vs. when building a bespoke solution is prudent Develop and execute against both short- and long-term roadmaps. Make effective tradeoffs that consider business priorities, user experience, and a sustainable technical foundation Who you are We’re looking for fullstack software engineers with experience building scalable products who have an eye for detail and are hap

reactaccounting
View job →
C
Cvshealth
📍 United States• Remote
11 days ago

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation

REMOTEkubernetesproject management
View job →
R
Ramp
📍 New York City• Full-time• From $10K/yr
1mo ago

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role The Security Engineering team helps make Ramp the most secure place for our customers to collect, manage, and put to work their business’ financial information Our work centers in three areas: Ramp builds products with an eye for security Ramp detects and responds to threats before they cause harm Security powers Ramp’s growth Check out our Engineering Blog for more on our tech stack, mission and values! What You’ll Do Drive our cloud security roadmap: review our cloud deployments to identify opportunities for improvement Design and build security-focused infrastructure primitives and integrate them into our existing products and development processes Lead remediation of prioritized issues across our technology stack Partner with infrastructure, data, and devops teams to design and deploy solutions that are inherently secure What You Need Minimum 5 years of experience building software Minimum 3 years of experience building in AWS (with Terraform) A strong sense of ownership: you need to drive projects from inception to scaling it in

pythonawsazure
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role As an Enterprise/Strategic Field Engineer (L5) , you'll be the technical cornerstone for Replit's largest and most strategic accounts. This is a hybrid role: high-impact pre-sales (closing complex technical evaluations) and post-sales (driving adoption, expansion, and retention). You'll own the end-to-end technical relationship—from pre-sales architecture discussions through multi-year expansion—ensuring our enterprise customers don't just use Replit, but become Replit-powered companies. You'll partner with Account Executives and Account Managers in a high-accountability Pod structure . This is not a reactive support role—this is a proactive, strategic technical leader who identifies blockers before they become problems, champions new use cases, and directly influences $5M+ in annual recurring revenue. In this role you will: Pre-Sales Strategic Technical Discovery: When you are pulled into complex deals, you join as the expert closer. You run deep discovery on their stack and constraints, then design the winning technical strategy. Proof of Value (POV) & Live Building: You build live, functional applications on the fly during executive meetings to prove immediate value and technical feasibility to VPs and C-suite stakeholders. Context & Connectivity (MCP): You write and deploy Model Context Protocol (MCP) servers to securely connect Replit Agents to customer-specific data, making Replit the central hub for their internal development. Enterprise Governance: You own the "Guardrails" mission. You configure workspace policies and AI governance templates that solve for data safety, compliance, and CISO approval. Infrastructure Strategy: You lead deep-dive reviews for Single-Tenant/VPC deployments, ens

javascriptpythonjava
View job →
R
Replit
📍 New York City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role As an Enterprise/Strategic Field Engineer (L5) , you'll be the technical cornerstone for Replit's largest and most strategic accounts. This is a hybrid role: high-impact pre-sales (closing complex technical evaluations) and post-sales (driving adoption, expansion, and retention). You'll own the end-to-end technical relationship—from pre-sales architecture discussions through multi-year expansion—ensuring our enterprise customers don't just use Replit, but become Replit-powered companies. You'll partner with Account Executives and Account Managers in a high-accountability Pod structure . This is not a reactive support role—this is a proactive, strategic technical leader who identifies blockers before they become problems, champions new use cases, and directly influences $5M+ in annual recurring revenue. In this role you will: Pre-Sales Strategic Technical Discovery: When you are pulled into complex deals, you join as the expert closer. You run deep discovery on their stack and constraints, then design the winning technical strategy. Proof of Value (POV) & Live Building: You build live, functional applications on the fly during executive meetings to prove immediate value and technical feasibility to VPs and C-suite stakeholders. Context & Connectivity (MCP): You write and deploy Model Context Protocol (MCP) servers to securely connect Replit Agents to customer-specific data, making Replit the central hub for their internal development. Enterprise Governance: You own the "Guardrails" mission. You configure workspace policies and AI governance templates that solve for data safety, compliance, and CISO approval. Infrastructure Strategy: You lead deep-dive reviews for Single-Tenant/VPC deployments, ens

javascriptpythonjava
View job →
PE
Private Employer
📍 Palo Alto• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Senior Identity Security Engineer on Palantir's Identity Security team, you will own the security posture of the identity infrastructure that Palantirians, customers, and services rely on every day. The Identity Security team is responsible for all identity types at Palantir - workforce, customer, workload, and agentic - giving you the rare ability to architect, threat model, and drive security outcomes across the full identity surface. You will help shape the technical direction for identity security at Palantir, reduce standing access, lead identity threat modeling, and contribute to the next generation of identity primitives including agent identity, JIT-native governance, and unified policy enforcement across workforce and customer IAM. As part of Palantir's best-in-class Information Security organization, you will research, architect, and scale solutions that help Palantir stay ahead of a dynamic identity threat landscape.

PE
Private Employer
📍 Washington• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Senior Identity Security Engineer on Palantir's Identity Security team, you will own the security posture of the identity infrastructure that Palantirians, customers, and services rely on every day. The Identity Security team is responsible for all identity types at Palantir - workforce, customer, workload, and agentic - giving you the rare ability to architect, threat model, and drive security outcomes across the full identity surface. You will help shape the technical direction for identity security at Palantir, reduce standing access, lead identity threat modeling, and contribute to the next generation of identity primitives including agent identity, JIT-native governance, and unified policy enforcement across workforce and customer IAM. As part of Palantir's best-in-class Information Security organization, you will research, architect, and scale solutions that help Palantir stay ahead of a dynamic identity threat landscape.

PE
Private Employer
📍 New York• Full-time• Hybrid
1mo ago

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role As a Senior Identity Security Engineer on Palantir's Identity Security team, you will own the security posture of the identity infrastructure that Palantirians, customers, and services rely on every day. The Identity Security team is responsible for all identity types at Palantir - workforce, customer, workload, and agentic - giving you the rare ability to architect, threat model, and drive security outcomes across the full identity surface. You will help shape the technical direction for identity security at Palantir, reduce standing access, lead identity threat modeling, and contribute to the next generation of identity primitives including agent identity, JIT-native governance, and unified policy enforcement across workforce and customer IAM. As part of Palantir's best-in-class Information Security organization, you will research, architect, and scale solutions that help Palantir stay ahead of a dynamic identity threat landscape.

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe's Developer & End-user Experience Platform (DEEP) organization empowers all of Stripe's products with a shared product platform that helps with rapidly delivering high-quality, cross-product experiences across our UI and API surfaces. It focuses on providing a consistent and scalable developer experience that any developer (both internal and external) can leverage to accelerate a merchant's ability to create value using Stripe. Team matching — exact team matching for one of the subteams will begin during final stages. Please note we may also consider you for different orgs based on your experience, location, etc. More information on our team matching process can be found here . What you’ll do We're looking for full-stack engineers who are interested in building software services and platforms that impact thousands of employees and millions of Stripe users, regardless of whether they're an end user, developer, or partner. Responsibilities Ensure our platforms are reliable, scalable, secure, and extensible Shape future-proof interfaces that are easy to build with Make effective tradeoffs that consider business priorities, user experience, and a sustainable technical foundation Help drive sound technical decision-making and lead technical conversations with other teams across Stripe Debug production issues across services and various levels of the tech stack Who you are We're looking for someone who meets the minimum requirements

restaigo
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

pythonawsrest
View job →
🔔

Get new lead software engineer infrastructure jobs by email

Daily job updates · Unsubscribe anytime