Jobiba hiring network

Lead Software Engineer Infrastructure Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead software engineer infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.

J
16 days ago

Backend Engineer (Senior Level) - SDE IV We're looking for a Senior Backend Engineer to lead the architecture and evolution of backend services that deploy and serve machine learning models in production. You'll work closely with ML Engineers, Platform, and Product teams to build scalable, reliable systems and drive technical direction across multiple teams. What You’ll Do Design and drive the long-term architecture of backend services for biometrics and ML model serving. Collaborate with core platform and backend teams on organization-wide architectural initiatives. Partner with business and engineering teams to design and deliver cross-cutting platform capabilities. Lead architectural reviews, mentor engineers, and promote engineering best practices. Build and maintain backend services for deploying and serving ML models Monitor service reliability, performance, and scalability in production Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 + Kubernetes Support real-time and batch inference workflows Contribute to CI/CD pipelines and deployment automation What We’re Looking For Strong expertise in backend development using Java and working knowledge of Python. Experience mentoring engineers and driving architectural decisions. Working knowledge of Python, especially for ML-related workflows Hands-on experience with AWS (e.g., DynamoDB, ECS, EC2, Redis, S3, SageMaker) Familiarity with Terraform or other infrastructure-as-code tools, and experience with CI/CD and production monitoring Experience with observability tools (Datadog, New Relic, etc.) Experience with containers and orchestration (Docker, ECS, etc.) Understanding of how ML models are deployed and served in production Experience with Kubernetes Nice to Have Experience with MLOps or ML platform engineering. Experience with asynchronous programming and event-driven systems. Jumio Values: IDEAL: Integrity, Diversity, Empowerment, Accountability, Leading Innovation Equal Opportunities :

REMOTEpythonjavaredis
View job →
N
12 days ago

We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea

awsazuredocker
View job →

The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease. We support products such as Bits AI , LLMObs and all our AI research . As an engineering manager for the Training & Serving team, you’ll join a new and fast growing team and organization. You will support building and scaling the team, define our technical vision and help shape the roadmap. Your team will lead the charge on multiple critical technical challenges: distributed training of foundation models, serving at scale, designing the user experience. You’ll work closely with sister teams in the AI platform organization ensuring a seamless AI development cycle. You’ll also partner with the Applied AI org and with Datadog infrastructure & tooling teams to build out systems from the ground up. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Manage and grow the Training & Serving team, directly managing 10+ engineers Define our technical roadmap in alignment with AI platform goals and the Applied AI team roadmap. Work with our core platform teams to tailor Datadog's storage, infrastructure and data pipelines to our needs Create a strong team culture aligned with our engineering standards and our customer focus Participate in hands-on work: Code reviews, design reviews and some coding Who You Are: Previous experience (1+ years) leading software engineering teams, as a tech lead or people manager Strong technician with a mix of backend, data engineer and infrastructure experience who is interested in remaining a hands-on leader Excellent leader with strong

restaigo
View job →
R
Roblox
📍 San Mateo• Full-time• From $243.3K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur

pythonawsgit
View job →

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro

pythonsqlpostgresql
View job →

Location Details: Pune, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team. Join our Team Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise! What you'll get to do... Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting

pythongitlinux
View job →
G
Gitlab
📍 United Kingdom; Remote, United States• Full-time• Remote
1mo ago

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a member of the Infrastructure Security Team within the Product Security Department , you will work with teams across GitLab to ensure that the components that comprise our public cloud infrastructure are built from the beginning with resiliency and set security expectations that our customers rely on to power their DevSecOps goals. As a Staff Security Engineer, you will serve as a technical lead across the topics the Infrastructure Security team owns, including our SaaS Platforms (e.g. GitLab Dedicated, Cells) and Self-Managed offerings. You will define the technical direction for how the team approac

REMOTEpythonawsazure
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Overview We are looking for a strong software and platform engineer to join our Production Engineering team in Bangalore as an individual contributor in FA ProductionEng APJ. This role will help build and operate internal platforms that improve how we provision, observe, govern, and troubleshoot engineering infrastructure at scale. The fleet management use cases that give teams a single place to understand and operate the test infrastructure. If you enjoy building internal platforms that remove friction, improve visibility, and make engineering teams faster and more effective, this role is for you. Why This Role Is Unique This is not a typical application development role.You will work on internal platforms that directly shape how engineering teams consume and manage shared infrastructure. The role spans platform engineering, workflow automation, observability, API-driven services, and infrastructure lifecycle management. The right candidate will work on systems such as: Self-serviceability workflows and lease-based testbed governance. Developer Platform dashboards and APIs used for triage, visibility, and product trend observation. Testbed and workflow orchestration across fleet management domains. Impact This role is a high-leverage engineering investment. The work will improve how engineering teams provision testbeds, understand failures, operate shared infrastructure, and move faster with less friction. Better

awskuberneteslinux
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxartificial intelligence
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxai
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Drive the hardware engineering lifecycle for Everpure’s Hyperscale Line of Business, shaping high-performance, energy-efficient storage architectures tailored directly to hyperscale partners . You will serve as the technical lead bridging customer requirements and internal engineering execution, ensuring custom hardware solutions meet exacting performance and total cost of ownership (TCO) benchmarks . Working closely with cross-functional teams in validation, diagnostics, manufacturing test, and operations, you will lead hardware qualification projects, solve critical escalations, and advance system robustness . This role gives you direct ownership over the hardware that powers next-generation, large-scale data infrastructure . WHAT YOU'LL DO Lead Hardware Qualification & System Hardening: Plan, execute, and automate comprehensive x86 hardware validation cycles—including electrical, signal integrity, and protocol testing—to guarantee system reliability for enterprise hyperscale deployments . Drive Cross-Functional Production Readiness: Partner with operations, manufacturing test, and software diagnostic teams to transition custom storage subsystems into full-scale production, establishing clear test coverage and automated validation workflows . Resolve High-Impact Technical Escalations: Perform root-cause analysis on complex hardware failures and field escalations, using advanced tools like high-speed o

pythonaisupply chain
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

pythonci/cdlinux
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking a highly technical Lead Release Engineer with a strong engineering foundation to lead the end-to-end release lifecycle of our storage products. You will serve as the bridge between Development, QA, and Product Management — ensuring that complex storage stacks are delivered with high quality and predictable cadences. Unlike traditional project-based release management, this role demands deep hands-on expertise across CI/CD orchestration, codeline management, system-level triaging, fleet operations, and the engineering rigor required for data-critical products. You will own the health of our release pipelines, lead triage war rooms, drive automation initiatives, and participate in on-call rotations to keep CI and test-orchestration infrastructure running reliably. You will also build developer-facing tooling, manage HW test fleet operations, and maintain high-quality integration workflows across our code lines. WHAT YOU'LL DO Release Orchestration: Own the end-to-end release process for storage software and firmware, from development to GA (General Availability). CI/CD Leadership: Design and build optimized pipelines and tools to scale code management and merge operations. Work closely with systems such as Jenkins, test frameworks, Premerge, Orchestrator, and related developer productivity tooling to keep the codeline healthy and actionable. Technical Triaging: Act as the primary technical poin

pythonawsci/cd
View job →
RS
12 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leader in full stack automation fabric solutions for mission-critical business processes. With the first SaaS-based composable automation platform specifically built for ERP, we believe in the transformative power of automation. Our unparalleled solutions empower you to orchestrate, manage and monitor your workflows across any application, service or server — in the cloud or on premises — with confidence and control. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Principal Engineer, Products & Platforms to join our Product engineering team to provide technical leadership across Redwood’s Workload Automation Platform, defining architecture, driving modernization, and influencing engineering strategy across multiple teams. You will be instrumental in the design, development, and enhancement of our platform, building high-quality, scalable, and secure software that powers enterprise data exchange for more than 1,000 customers worldwide. As a Principal Engineer, you will: Technical Leadership & Architecture: Define the path forward for complex engineering problems, establish best practices, lead design review and technology decisions, and mentor the team to excel, while driving the architecture, security, compliance, and observability of our Java/Spring Boot microservices. Platform and Infrastructure Ownership: Drive the deep understanding, architecture, and evolution of our core platform and infrastructure, focusing on resilience, communication between components, and scalability. Cross-Team Collaboration: Facilitate and drive cross-team collaboration with Product, QA, and other engineering groups to ensure end-to-end alignment and successful product delivery. AI Integration:

javaawskubernetes
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security a

airustexcel
View job →
🔔

Get new lead software engineer infrastructure jobs by email

Daily job updates · Unsubscribe anytime