About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
Jobiba hiring network
Lead Software Engineer Infrastructure Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead software engineer infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,
About the team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the role: We're seeking a Data Engineer to take the lead in building our data pipelines and core tables for OpenAI. These pipelines are crucial for powering analyses, safety systems that guide business decisions, product growth, and prevent bad actors. If you're passionate about working with data and are eager to create solutions with significant impact, we'd love to hear from you. This role also provides the opportunity to collaborate closely with the researchers behind ChatGPT and help them train new models to deliver to users. As we continue our rapid growth, we value data-driven insights, and your contributions will play a pivotal role in our trajectory. Join us in shaping the future of OpenAI! In this role, you will: Design, build and manage our data pipelines, ensuring all user event data is seamlessly integrated into our data warehouse. Develop canonical datasets to track key product metrics including user growth, engagement, and revenue. Work collaboratively with various teams, including, Infrastructure, Data Science, Product, Marketing, Finance, and Research to understand their data needs and provide solutions. Implement robust and fault-tolerant systems for data ingestion and processing. Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear. Ensure the security, integrity, and compliance of data according to industry and company standards. You might thrive in this role if you: Have 3+ years of experience as a data engineer and 8+ years of any software engineering experience(including data engineering). Proficiency in at least one programming language commonl
About Ubiquiti At Ubiquiti Inc., we create technology platforms for Businesses, Smart Homes, and Internet Service Providers, driven by our goal to connect everyone, everywhere. To date, Ubiquiti has shipped over 100 million devices worldwide, from ISP networking products to next generation of IT solutions. Our growth is made possible by the dedicated team of hundreds behind the scenes. From software developers and product managers to designers and strategists, Team UI is driven to achieve our common goal: Rethinking IT. At Ubiquiti, you’ll heighten your potential and broaden your horizons - all while shaping the future of connectivity. Responsibilities Lead hardware circuit design, schematic capture, and layout review for next-generation, high-density, and high-throughput networking platforms. Evaluate and integrate advanced switching architectures, management subsystems, and cutting-edge high-speed interconnect technologies. Drive hardware architecture design, system bring-up, high-speed signal integrity (SI) validation, and root-cause failure analysis. Partner with mechanical, thermal, and power engineering teams to address challenges related to high power density, thermal dissipation, and system-level reliability. Own BOM structure and support factory deployment to ensure seamless transition of high-layer-count PCBAs from NPI to mass production. Collaborate with cross-functional software, firmware, QA, and compliance teams throughout the entire product lifecycle. Q ualifications Bachelor’s degree or above in Electrical Engineering or a related discipline. 5+ years of hands-on experience in high-complexity system-level hardware design, ideally focused on enterprise-grade networking or high-performance infrastructure equipment. Deep technical understanding of high-speed Ethernet design, high-speed differential signals (advanced SerDes, PCIe, multi-gigabit/ultra-high-speed interfaces), and high-density PCB design rules. Practical experience with complex power d
Senior Container Security Engineer – CVE Remediation & Image Hardening About the Role We are looking for a hands-on Senior Container Security Engineer to lead vulnerability remediation and image hardening across Linux-based container environments. This role focuses on deep operating system and container security engineering rather than simple vulnerability scanning. You will analyze, remediate, rebuild, harden, and continuously optimize container images used in modern cloud-native platforms. You will work closely with platform engineering, DevOps, infrastructure, and security teams to build automated remediation pipelines, reduce the attack surface, and deliver production-ready hardened images. What You’ll Do - Own end-to-end CVE remediation across Linux-based container images. - Analyze vulnerabilities across OS packages, libraries, runtimes, and dependencies. - Patch, rebuild, validate, and maintain hardened container images at scale. - Reduce attack surface by removing unnecessary packages, binaries, services, and dependencies. - Build and scale automated remediation pipelines for continuous image patching. - Improve image security posture while minimizing operational disruption. - Generate, validate, and maintain SBOMs to support supply chain visibility and compliance. - Integrate remediation workflows into CI/CD and GitOps pipelines. - Optimize image size, startup performance, and operational efficiency. - Research emerging Linux, container, Kubernetes, and software supply chain threats. - Troubleshoot complex dependency, package compatibility, and runtime security issues. - Help define internal standards for hardened images and secure software delivery. What You Bring - 5+ years of experience in Linux systems engineering, platform engineering, DevSecOps, security engineering, or SRE. - Deep understanding of Linux distributions (Debian, Ubuntu, Alpine, RHEL). - Strong hands-on experience with Docker, Kubernetes, and
Location: San Francisco, CA (Hybrid) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role We're seeking an experienced Senior Optimization Engineer to join our Data Science Team. In this role, you will lead the design, development, and deployment of optimization models that power our software platform across applications including electricity markets, renewable energy, and battery energy storage systems. You will be responsible for developing production-grade optimization engines that solve complex operational and planning problems at scale. This role requires deep expertise in mathematical optimization, strong software engineering skills in Python, and experience building optimization models that integrate with production systems. The ideal candidate has significant experience in the energy industry, particularly electricity markets and battery storage optimization. This position emphasizes technical leadership, ownership of complex optimization projects, and collaboration across engineering, product, and commercial teams to deliver high-impact optimization solutions. Key Res
About the Team DoorDash Labs is an independent team within DoorDash. We explore robotics and automation to transform last mile logistics in the long term. If you have a passion for applying robotics solutions in a service used by millions of people, then we want to talk to you! About the Role We are hiring a Firmware Validation & Integration Engineer for our autonomy software team. This is a critical role to build robust and scalable validation for our firmware and systems to ensure reliability at every level. In this role, you will work with our electrical, firmware, and autonomy engineers to build the infrastructure and test suites required to validate the system. This includes designing and implementing our Hardware-in-the-Loop (HIL) simulation environments and automation frameworks from the ground up. You will report to the Autonomy Platform Lead on our Autonomy Platform Team at DoorDash Labs. We expect this role to be hybrid with some time in-office and some time remote. You’re excited about this opportunity because you will… Play an integral role on a small and focused team. Design and build Hardware-in-the-Loop (HIL) systems to simulate vehicle dynamics and sensor data for comprehensive firmware and system-level validation. Develop automated test infrastructure and software tools to exercise multiple embedded platforms throughout our robot system. Interface many layers of our control system including vehicle controls, power management, and motion control to ensure seamless system integration. Implement low-level test sequences and validation algorithms to safely stress-test vehicle components such as batteries, drive-train, and thermal management devices. Collaborate with cross-functional teams to identify edge cases and hardware-software corner cases that impact vehicle safety and performance. We’re excited about you because… BS/MS degree in Computer Science, Robotics, Electrical Engineering, or related technical field. 5+ years of experience in validati
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are looking for an experienced Mechanical Engineer with 7+ years of experience in design of IT hardware from chip/package to system levels. You’ll work alongside experts in thermal, mechanical, electrical, software, and systems engineering to support the design, analysis, and validation of mechanical and thermal systems that ensure the reliability, efficiency, and longevity of mission-critical hardware. This position requires strong analytical skills, hands-on testing experience, and the ability to work in a fast-paced, cross-disciplinary environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead mechanical design for AI supercomputer product in the data center application Collaborate with the cross functional team to design and optimize thermal solutions for data center hardware, including chips, power modules, and system-level cooling architectures Collaborate with cross-functional teams to integrate thermal management strategies into hardware design, from concept to mass production Design and validate mechanical systems, including chassis, enclosures, cooling systems, and high-power connections, ensuring alignment with performance and reliability standards. Perform 3D modeling, FEA, tolerance analysis, and prototyping, ensuring manufacturability and a
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is seeking an experienced Senior Adobe Experience Cloud Engineer with a deep understanding of Adobe’s tech stack to join our growing team.The position will play a crucial role in designing, developing, and maintaining solutions that leverage Adobe technologies to meet the unique business needs of our business partners. You must be collaborative and able to build trusted partnerships with extended teams. A successful candidate will have the ability to balance priorities and collaborate with cross functional teams while delivering within an agile delivery framework and supervising key performance indicators. Responsibilities Solution Design & Development: Lead the development and implementation of custom solutions and integrations within Adobe Experience Cloud, including Adobe Experience Cloud solutions, including Content Management, Assets, Multi-Site-Management, and Cloud manager Architecture & Scalability: Architect and build scalable, high-performance systems and applications that meet business requirements and technical specifications. Integration & Optimization: Develop and maintain integrations between Adobe Experience Cloud products and other internal or third-party systems. Optimize existing systems for performance and reliability. Collaboration: Work closely with product managers, solution architects, and other stakeholders to gather requirements, define project scopes, and deliver high-quality software solutions. Agile Development:
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About Okta for AI Agents Okta secures access for 20,000 organizations and billions of users. Okta for AI Agents extends that work to the agentic shift. Deploying an AI agent is not like deploying traditional software. You are putting professional work output into production, and it needs deep integration, continuous tuning, and change management. Every agent needs an identity, a scope, an audit trail, and a way to be shut down when it goes wrong. Most enterprises have not built this yet. We are. We hire builders who see the cracks in enterprise agent identity that everyone else has learned to live with. The Role You are the most senior technical field authority for agent identity at Okta. Where a Senior FDE owns the outcome inside one account, you own the patterns that every account and every FDE inherits. You take the hardest and most strategic deployments yourself, set the reference architecture the team builds from, and turn what the field learns into the direction the product takes. You still write code. You also multiply the people around you, and you are the person product and engineering leadership call when an agent identity problem has no precedent. Responsibilities Own the reference architecture. Define the canonical agent identity, delegation, audit, and kill-switch patterns that Senior FDEs deploy across the portfolio, and keep them current as the standards and the product move. Lead the hardest accounts. Personally own the most strategic, regul
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About Okta for AI Agents Okta secures access for 20,000 organizations and billions of users. Okta for AI Agents extends that work to the agentic shift. Deploying an AI agent is not like deploying traditional software. You are putting professional work output into production, and it needs deep integration, continuous tuning, and change management. Every agent needs an identity, a scope, an audit trail, and a way to be shut down when it goes wrong. Most enterprises have not built this yet. We are. We hire builders who see the cracks in enterprise agent identity that everyone else has learned to live with. The Role You are the most senior technical field authority for agent identity at Okta. Where a Senior FDE owns the outcome inside one account, you own the patterns that every account and every FDE inherits. You take the hardest and most strategic deployments yourself, set the reference architecture the team builds from, and turn what the field learns into the direction the product takes. You still write code. You also multiply the people around you, and you are the person product and engineering leadership call when an agent identity problem has no precedent. Responsibilities Own the reference architecture. Define the canonical agent identity, delegation, audit, and kill-switch patterns that Senior FDEs deploy across the portfolio, and keep them current as the standards and the product move. Lead the hardest accounts. Personally own the most strategic, regul
Staff Engineer, Revenue and Financial Management Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe’s Revenue & Financial Management group is building new products that expand the scope of problems we tackle beyond payments into Revenue Management, Financial Operations and analytics. Right now this includes products like Billing, Invoicing, Revenue Recognition, Reconciliation and the underlying platform for batch and real time processing of large scale financial data to power these solutions. These solutions are going to be key pillars for Stripe’s growing SaaS business and a major revenue stream. What you’ll do We’re looking for a Staff Engineer that will help architect and design this system from ground up. You will need to set the technical direction across a variety of projects and initiatives while also mentoring and growing others on the team. Responsibilities Scope and lead large technical projects that are the foundational pillars for Financial Data Management Infrastructure Scrutinize and reason clearly about the technology and architecture choices we make in building these products. In many cases, you will be the decider of these decisions Directly contribute to core interface design and write code. Serve as a role model for how great software should be written for Stripe as a whole Arbitrate critical decisions correctly that fully consider software best practices, Stripe system realities, and numerous stakeholders’ preferences and concerns Advise Stripe’s leadership tea
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi
What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. We’re hiring a Manufacturing Reliability Engineer to own production test for our robots at our contract manufacturer: you’ll design and run robust end-to-end test protocols, provision fleets of robots for production, and own the KPIs that define production quality. This role is based in Austin, TX. However, the position will require 50% travel to the Milwaukee, WI area and requires close collaboration across software, hardware, operations, and product engineering teams. Key Responsibilities End-to-end test process ownership. Create, validate, and maintain production test protocols and gating criteria from incoming inspection through final test and shipment. Provisioning of bots. Design and operate provisioning flows (imaging, firmware deployment, configuration, validation) and the tooling/fixtures needed to provision and handoff robots for production. KPIs and continuous improvement. Own key production metrics — First Pass Yield (FPY), cycle time, and test coverage — and drive continuous improvements to meet throughput and quality targets. Test automation & infrastructure. Architect, implement, and maintain automated test frameworks, harnesses, and test rigs used at the CM site. Ensure tests are stable, fast, and provide actionable failure data. Cross-functional escalation & RCA. Lead root-cause analysis for field and production failures; coordinate corrective actions with design, firmware, and CM engineering to close quality loops. On-site production leadership. Be the onsite technical authority at the contract manufacturer: train operators, debug failures on the line, and continuously refine processes with CM partners. What Success Looks Like Improved FPY and reduced rework rates across production builds. Reduced per
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
Get new lead software engineer infrastructure jobs by email
Daily job updates · Unsubscribe anytime