Jobiba hiring network

Lead Software Engineer Infrastructure Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead software engineer infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.

CI
Couchbase, Inc.
📍 Bengaluru• Full-time
15 days ago

Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! Lead Software Engineer - Storage As a key contributing member of the storage development team, you will be responsible for enhancing the highly scalable and performant storage engines used by both Couchbase Server and Couchbase Capella. You will be developing a highly-available and concurrent enterprise-grade system software. Most of all, you will be able to celebrate the wins by experiencing the direct result of your hard work from our customers’ success stories. The ideal candidate will have a strong technical background, excellent communication skills, and proactive problem-solving skills. The innovative work storage team does has been widely recognized by the industry. The following publications in VLDB conferences reflect the storage work at Couchbase Nitro: A Fast, Scalable In-Memory Storage Engine for NoS

javasqlaws
View job →
NR
New Relic
📍 India• Full-time
1mo ago

We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your opportunity At New Relic, we provide our customers real-time insights, so they can innovate faster. Our software delivers insightful observability tools across different technologies and distributed systems, enabling software engineering teams to quickly identify, understand and tackle issues, analyze performance and get the most of their software and infrastructure. The Infrastructure product organization develops New Relic infrastructure instrumentation agents, next generation data processing and management services, vulnerability management, and security testing capabilities for on-prem and cloud customers. We work with data at a scale using a diverse tech stack (Go, Java, JavaScript, React GraphQL, Kubernetes, many public cloud web services, and more). As a senior backend engineer, you will help us build and extend next generation solutions such as a control plane for customers to manage their data pipelines at scale. New Relic is looking for engineers who are interested in building a brand-new observability experience. This high-impact engineering position is a phenomenal opportunity to own and build a set of next generation services and capabilities for the company. We are searching for a motivated engineer who is ready for a career-defining role in their next opportunity. We look forward to talking with you! What you'll do ● Design, Build, maintain, and scale back-end services and their support tools. ● Participate in architectural definitions with a high degr

javascriptjavareact
View job →

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Software Engineer - Java Are you an experienced software engineer who enjoys solving complex technical challenges and helping others build better solutions? As a Lead Software Engineer, you will design, build, and operate secure, scalable, and resilient technology that supports important payment services. You will combine hands-on development with technical leadership and work closely with teams across product, operations, architecture, infrastructure, risk, and customer-facing functions. Mastercard Payment Services: In Mastercard Payment Services, you will contribute to critical European payment clearing and settlement services. This includes real-time and batch clearing, instant payments, participant-facing services, liquidity and settlement capabilities, reporting, and related payment infrastructure. You will help ensure these services are reliable, secure, and resilient by applying strong engineering practices, automation, production ownership, and continuous improvement. What you will do: • Design, build, integrate, and improve secure, scalable applications and services. • Take ownership of complex service issues and coordinate resolution across teams. • Work with product partners and stakeholders on priorities, trade-offs, and technical roadmaps. • Improve delivery quality through a

javareactdocker
View job →
M
10 days ago

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Software Engineer - Java Are you an experienced software engineer who enjoys solving complex technical challenges and helping others build better solutions? As a Lead Software Engineer, you will design, build, and operate secure, scalable, and resilient technology that supports important payment services. You will combine hands-on development with technical leadership and work closely with teams across product, operations, architecture, infrastructure, risk, and customer-facing functions. Mastercard Payment Services: In Mastercard Payment Services, you will contribute to critical European payment clearing and settlement services. This includes real-time and batch clearing, instant payments, participant-facing services, liquidity and settlement capabilities, reporting, and related payment infrastructure. You will help ensure these services are reliable, secure, and resilient by applying strong engineering practices, automation, production ownership, and continuous improvement. What you will do: • Design, build, integrate, and improve secure, scalable applications and services. • Take ownership of complex service issues and coordinate resolution across teams. • Work with product partners and stakeholders on priorities, trade-offs, and technical roadmaps. • Improve delivery quality through a

javareactdocker
View job →
HI
HP IQ
📍 San Francisco• Full-time• $179K – $252K/yr
15 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Lead Software Engineer, Cloud Services, you will create scalable, reliable backend systems that support HP IQ's mission to transform the way people work. We value how we work as much as what we deliver. Our journey of continuous learning and evolution requires a flexible and resilient services platform that fosters innovation, experimentation, and adaptation. To enable this, we focus on designs and tools rooted in strong engineering principles like abstraction, composition, virtualization, automation, and iterative development cycles. What You Might Do Lead technical strategy and execution across development and infrastructure, designing and implementing scalable, secure cloud-native systems while establishing architecture standards and engineering best practices. Own end-to-end delivery of platform and cloud services, including infrastructure buildout, application architecture, APIs, deployment pipelines, observability, reliability, and performance, while mentoring engineers and driving cross-functional alignment. Work with modern containerization and cloud technologies, includi

javaredisdocker
View job →
HI
HP IQ
📍 San Francisco• Full-time• $190K – $270K/yr
15 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The AI team is building cutting-edge solutions that bring the power of AI directly to edge devices while seamlessly integrating with cloud infrastructure. We are looking for a Lead Software Engineer to design and develop high-performance, scalable services to support AI workloads across edge and cloud environments. What You Might Do Design, build, and maintain services that power AI-driven applications, ensuring scalability and performance. Develop APIs and microservices that facilitate seamless integration between cloud-based AI models and edge devices. Optimize data pipelines and storage solutions for real-time AI inference and processing. Implement security and privacy best practices for distributed AI systems. Work closely with AI researchers, infrastructure engineers, and frontend developers to deliver end-to-end AI-driven solutions. Build and optimize an agent orchestration runtime that enables tool use, memory management, and multi-step reasoning across LLMs, APIs, and edge-connected systems. Develop robust logging, monitoring, and alerting systems to ensure system reliability.

pythonjavasql
View job →
CH
15 days ago

Opportunity Overview: We are seeking a Lead Software Engineer to join our Integrations team. In this role, you will be designing, developing, and scaling highly available healthcare integration systems supporting prior authorization workflows across providers, payers, and delegated entities. You'll direct a fast-paced, autonomous,agile team of software engineers in the design, development, and operational support of a growing enterprise integration platform. This is an opportunity to drive technical excellence at the intersection of healthcare interoperability and modern distributed systems. What you’ll do: Technical Leadership: Provide technical leadership across architecture, system design, platform scalability, reliability, and operational excellence. Platform Engineering: Design and build scalable, resilient, and high-performing systems that support critical business workflows and enterprise integrations. Integration Solutions: Lead the development and maintenance of secure integrations with internal and external platforms, partners, and third-party systems. Cloud & Automation: Drive cloud infrastructure, deployment automation, and software delivery practices that enable reliable and efficient releases. Distributed Systems: Design and support event-driven and distributed architectures that enable scalable and fault-tolerant processing. Operational Excellence: Establish monitoring, observability, and incident response practices to ensure system reliability, performance, and availability. Quality Engineering: Champion automated testing, quality assurance, and engineering best practices throughout the software development lifecycle. Production Support: Lead the resolution of complex production issues and drive continuous improvement in platform stability and operational efficiency. Cross-Functional Collaboration: Partner with product, operations, data, security, and business stakeholders to deliver solutions aligned with organizational goals. Agile Delive

javaawsdocker
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr

javavueaws
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is building its own GPU infrastructure for large-scale inference. As we move into large scale, high-density NVIDIA systems, the hardest failures are intermittent, cross-layer, and difficult to prove: RoCE congestion, InfiniBand stalls, ECN/DCQCN mis-tuning, bad optics, RNIC issues, host kernel stalls, GPU driver problems, and workload symptoms that look like network problems, but are not. We are hiring a Lead Software Engineer to build a first-class observability and root-cause analysis system for GPU fabrics. This is a hard distributed systems problem, not a dashboarding problem. The system will collect high-volume signals from switches, hosts, active probes, and inference services; reduce and correlate them in real time; understand topology and service ownership; and produce actionable diagnosis while an incident is still unfolding. This role sits at the boundary between networking and inference software. RDMA data paths, GPUDirect transfers, prefill/decode disaggregation, KV cache movement, request routing, and workload backpressure can all create fabric symptoms or hide real fabric failures. The goal is to tell an operator, quickly and with evidence, whether an incident is caused by the fabric, host, NIC, GPU, RDMA path, scheduler, or serving layer — and what to do next. EXAMPLE INITIATIVES Real-time telemetry engine — Build the ingestion, reduction, storage, and query path for high-cardinality fab

kubernetesmachine learningai
View job →
B
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg

pythonawsgcp
View job →
D
Dropbox
📍 Poland• Full-time• Remote
1mo ago

Role Description As an Infrastructure Engineer, your role will be crucial in shaping and constructing the robust systems that not only support our current flagship products but also lay the groundwork for the next wave of engineering innovations. From optimizing user experiences across various projects to ensuring seamless scalability and data integrity, you'll be at the forefront of shaping the technological backbone of our platform. Collaborating closely with cross-functional teams, you'll leverage your expertise to tackle audacious challenges and push the boundaries of what's possible. Your contributions will directly impact millions of users, as every line of code you write furthers our mission to revolutionize the way people work and collaborate. Join us in redefining the future, where your passion for building scalable, reliable systems will drive meaningful change on a global scale. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Build infrastructure capable of managing metadata for hundreds of billions of files, handling hundreds of petabytes of user data, and facilitating millions of concurrent connections. Lead the expansion of Dropbox's function as the data-fabric, connecting hundreds of millions of applications, devices, and services globally, while also driving initiatives to enhance interoperability and adaptability across diverse ecosystems. M easur e and optimiz e Dropbox's analytics platform to maintain its status as one of the most advanced in the industry for extracting meaningful insights from vast data volumes. Collaborat e with cross-functional teams to innovate and implement solutions that enhance the performance, reliability, and security of Dropbox's infrastructure, ensuring a seamless experience for users worldwide. Proactively identify new opportunities and drive imp

REMOTEpythonjavagit
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that our model's training environment is as realistic as possible. This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness. About the Role We're looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You'll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments. This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality. In this role, you will Design, build, and evolve the integration between the Codex harness that powers OpenAI's products and research training infrastructure used for training GPT-next Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance Bu

pythonawsrest
View job →
E
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a Software Engineer on the DX Security Components team , you will architect the backend services that safeguard Pure’s core authentication and remote access infrastructure. You’ll act as a security-minded engineer within a high-impact platform team, ensuring our cloud offerings and customer appliances remain auditable and resilient. Collaborating across the DX organization, you will bridge the gap between robust security protocols and seamless developer integration to protect our global customer base. WHAT YOU’LL DO Engineer Security Infrastructure: Design and operate high-availability backend services that manage authentication, authorization, and certificate lifecycles to ensure secure access across all Pure1 cloud and appliance environments. Drive End-to-End Ownership: Lead the full service lifecycle—from initial architectural design and threat modeling (STRIDE) to deployment, observability, and long-term cost efficiency. Champion Secure Integration: Partner with Security Governance and product teams to streamline remote-access flows, translating complex security requirements into pragmatic, automated workflows for other engineering squads. Ensure System Resilience: Maintain the integrity of security-sensitive systems by participating in a global follow-the-sun on-call rotation, performing root-cause analysis, and hardening infrastructure against emerging threats. Automate Trust: Evolve Infrastructure as Cod

pythonawsci/cd
View job →
A
Amplitude
📍 San Francisco• Full-time
15 days ago

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Winter 2026 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. About the Team DevX owns engineering velocity at Amplitude: build systems, CI/CD, developer environments, and internal tooling. We're building a software factory: automated workflows that remove manual bottlenecks from how engineers ship code. The Role We're looking for a Staff Software Engineer – DevX (Hybrid – San Francisco) who bridges infrastructure and application thinking and can accelerate how the whole team develops, tests, and ships in the cloud. You'll set architecture for our developer platform, lead our software factory work, and push our development model toward cloud-first workflows. This is a high-leverage, low-oversight role. You'll own initiatives end to end, from an ambiguous problem to production, and set technical direction for a foundational team. What You'll Do Cloud development platform: Unify and scale our existing loca

ci/cdgitai
View job →

Staff Software Engineer - UI Foundation at Amplitude (View all jobs) San Francisco Bay Area Hybrid Amplitude is the leading AI-first digital analytics platform, helping over 4,300 customers—including Atlassian, Burger King, NBCUniversal, Square, and Under Armour—build better products and digital experiences. With powerful AI Agents embedded across our platform, teams can analyze, test, and optimize user experiences faster than ever. Ranked #1 across multiple categories in G2’s Fall 2025 Report, Amplitude is the best-in-class solution for product, data, and marketing teams. Learn more at amplitude.com . As an organization, we deliver for our customers by living our values. We operate from a place of humility, take ownership of problems and successes, approach challenges with a growth mindset, and put our customers at the center of everything we do. Amplitude’s Commitment to Diversity Equity & Inclusion (DEI): Amplitude believes that diversity enables the creation of better products, improves the ability to solve complex problems, and drives more powerful solutions. We strive to create an environment of inclusion—one focused on psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive. The UI Foundation team owns the frontend infrastructure that every product engineer at Amplitude builds on — architecture, performance, developer experience, tooling, and build systems. We're the team that makes shipping fast, reliable frontend code possible at scale, and we're increasingly focused on how AI agents can autonomously handle engineering workflows that used to require constant human intervention. The Role We're looking for a Staff Software Engineer who thinks like an infrastructure engineer but has deep fluency in frontend systems. You'll drive architectural decisions for our frontend platform, push performance and DevX forward, and lead the build-out of our "software factory" — a set of autonomous, AI-agent-driven engine

reactci/cdgit
View job →
🔔

Get new lead software engineer infrastructure jobs by email

Daily job updates · Unsubscribe anytime