Jobs in United States

Lead Software Engineer Infrastructure in United States

2,434 active opportunities · Updated October 2026

Explore current lead software engineer infrastructure jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

51/100

steady · 543 related jobs

Hiring trend

-76.9%

Job postings compared with the previous 30 days

Remote options

15.1%

Share of matching jobs listed as remote

Typical salary

$177.2K – $177.2K/yr

Based on 29 salary observations

G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects. Job Summary We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines. The Team The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers. The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale. As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems

PythonKubernetesCI/CDGit
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in artificial intelligence compute. We are developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and support the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a family of companies responsible for some of the world’s most transformative technologies. Together, we share a bold vision to enable advanced artificial intelligence and ensure its benefits are accessible to everyone. Graphcore brings together AI researchers, silicon designers, software engineers and systems architects to solve complex technical challenges and deliver innovative computing solutions. Job Summary The Principal Electrical Engineer will be a technical authority within Data Center Engineering, leading the architecture and delivery of safe, resilient and scalable electrical infrastructure for high-density AI computing environments. Working with internal teams, data center developers, utilities, consultants and equipment partners, this role will guide projects from early technical studies through design, construction, commissioning, operation and lifecycle improvement. The successful candidate must reside in, or be willing to relocate to, Austin, Texas. Approximately 10% travel may be required. The Team The Data Center Engineering team is responsible for defining and enabling the infrastructure needed to deploy and operate Graphcore’s computing systems at scale. The team works across electrical, mechanical, thermal, controls, systems and operational disciplines, collaborating with external engineering and construction partners to deliver reliable, efficient and maintainable data center environments. Responsibilities and Duties Act as the technical authority for electrical engineering across data center infrastructure projects, from the utility or on-site power source through to the IT rack. Lead electrical archit

AIExcelSEMAutocad
I
📍 Oregon, Hillsboro, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: Join an enthusiastic team of engineers in Intel's Networking Solutions Group (NSG) focused on enabling next generation of programmable Infrastructure Processing Units (IPUs) with our lead customers as part of the Customer Experience Support (CES) organization. Intel brings decades of leadership in networking, virtualization, packet processing, storage, and security to a new class of IPU products that accelerate host networking functions and support emerging use cases such as security, virtualization, storage, load balancing, and data path optimization. Working closely with major cloud service providers and Intel development teams, you will help deliver customized IPU based solutions that enhance isolation, security, performance, storage and system management for our customers. A big part of the day-to-day job is to help customers manage feature request processes, enable solutions, and debug issues. Projects and responsibilities include but are not limited to: • Gain our customers' trust, understand their needs, and build POCs to meet them. Work closely with internal and external partners to understand use cases and requirements. • Be the go-to technical resource for customers building complex Datacenters, AI infrastructure as well as helping them understand performance characteristics for solutions. • Prepare and deliver technical content to customers including presentations, workshops, etc. • Contribute across the full IPU lifecycle, including board and platform bring up, low-level device initialization, OS driver and kernel configuration, system management, feature enablement, use case testing, debugging, and verification. • Defines systems implementation and integration solutions and plans to ensure optimum performance and reliability across hardware, firmware and software w

DockerGitLinuxAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will lead the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with cross-functional teams, you will contribute to the design and implementation of SLT hardware, software, automation, and characterization solutions that support silicon bring-up, yield learning, manufacturing readiness, and production deployment. The ideal candidate will possess strong technical depth in semiconductor test and validation, a passion for solving complex engineering challenges, and a strong focus on product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Lead development and deployment of SLT hardware and software solutions supporting silicon bring-up, charac

PythonAISEMHR
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

PythonAIGoDevOps
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The Applied team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the role: We're seeking a Data Engineer to take the lead in building our data pipelines and core tables for OpenAI. These pipelines are crucial for powering analyses, safety systems that guide business decisions, product growth, and prevent bad actors. If you're passionate about working with data and are eager to create solutions with significant impact, we'd love to hear from you. This role also provides the opportunity to collaborate closely with the researchers behind ChatGPT and help them train new models to deliver to users. As we continue our rapid growth, we value data-driven insights, and your contributions will play a pivotal role in our trajectory. Join us in shaping the future of OpenAI! In this role, you will: Design, build and manage our data pipelines, ensuring all user event data is seamlessly integrated into our data warehouse. Develop canonical datasets to track key product metrics including user growth, engagement, and revenue. Work collaboratively with various teams, including, Infrastructure, Data Science, Product, Marketing, Finance, and Research to understand their data needs and provide solutions. Implement robust and fault-tolerant systems for data ingestion and processing. Participate in data architecture and engineering decisions, bringing your strong experience and knowledge to bear. Ensure the security, integrity, and compliance of data according to industry and company standards. You might thrive in this role if you: Have 3+ years of experience as a data engineer and 8+ years of any software engineering experience(including data engineering). Proficiency in at least one programming language commonl

PythonJavaAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We are looking for an experienced Mechanical Engineer with 7+ years of experience in design of IT hardware from chip/package to system levels. You’ll work alongside experts in thermal, mechanical, electrical, software, and systems engineering to support the design, analysis, and validation of mechanical and thermal systems that ensure the reliability, efficiency, and longevity of mission-critical hardware. This position requires strong analytical skills, hands-on testing experience, and the ability to work in a fast-paced, cross-disciplinary environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead mechanical design for AI supercomputer product in the data center application Collaborate with the cross functional team to design and optimize thermal solutions for data center hardware, including chips, power modules, and system-level cooling architectures Collaborate with cross-functional teams to integrate thermal management strategies into hardware design, from concept to mass production Design and validate mechanical systems, including chassis, enclosures, cooling systems, and high-power connections, ensuring alignment with performance and reliability standards. Perform 3D modeling, FEA, tolerance analysis, and prototyping, ensuring manufacturability and a

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi

AWSRestAIRust
DR
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. We’re hiring a Manufacturing Reliability Engineer to own production test for our robots at our contract manufacturer: you’ll design and run robust end-to-end test protocols, provision fleets of robots for production, and own the KPIs that define production quality. This role is based in Austin, TX. However, the position will require 50% travel to the Milwaukee, WI area and requires close collaboration across software, hardware, operations, and product engineering teams. Key Responsibilities End-to-end test process ownership. Create, validate, and maintain production test protocols and gating criteria from incoming inspection through final test and shipment. Provisioning of bots. Design and operate provisioning flows (imaging, firmware deployment, configuration, validation) and the tooling/fixtures needed to provision and handoff robots for production. KPIs and continuous improvement. Own key production metrics — First Pass Yield (FPY), cycle time, and test coverage — and drive continuous improvements to meet throughput and quality targets. Test automation & infrastructure. Architect, implement, and maintain automated test frameworks, harnesses, and test rigs used at the CM site. Ensure tests are stable, fast, and provide actionable failure data. Cross-functional escalation & RCA. Lead root-cause analysis for field and production failures; coordinate corrective actions with design, firmware, and CM engineering to close quality loops. On-site production leadership. Be the onsite technical authority at the contract manufacturer: train operators, debug failures on the line, and continuously refine processes with CM partners. What Success Looks Like Improved FPY and reduced rework rates across production builds. Reduced per

PythonAIExcelHR
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $244K/yr

Quick readStrong listing-quality and freshness signals

Datadog’s Cloud Networks team designs, builds, and maintains the production network infrastructure that powers everything built on top of our platform across AWS, GCP, Azure, and beyond. In this role, you’ll set technical direction for how we scale our multi-region, multi-cloud network footprint while keeping reliability and performance high. You’ll partner closely with internal teams and Cloud Service Providers to troubleshoot complex connectivity issues, integrate new networking capabilities, and improve the foundations our engineers and customers rely on. This is a high-impact opportunity to drive meaningful improvements in scale, resiliency, and cost efficiency. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design, build, and operate cloud network infrastructure across AWS, GCP, Azure, and Neoclouds in a multi-region environment. Own connectivity between clouds, customers, and developers—ensuring scalable, secure, and reliable network paths. Set clear technical direction for expanding data centers and evolving the network while maintaining stability and performance. Improve cross-site and cross-region connectivity patterns to support Datadog’s growing platform needs. Lead deep investigations into latency, packet loss, and connectivity failures – from pcap and path analysis through to escalations with cloud providers that may originate from customer support Identify and deliver network-related efficiency and cost-saving opportunities that positively impact business health. Who You Are: You have deep networking expertise. You understand BGP, route policies, path selection, prefix advertisement, and what breaks in large-scale networking. You have substantial experience designing, building, and evolving large-scale Software-Defined Networks—inclu

AWSAzureGCPAI
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%

From $132.4K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Are you passionate about building impactful products for sales and finance teams? Come join the IT Enterprise Systems team at Pinterest where you will be responsible for advancing our sales and marketing systems. What you’ll do: Design, build, and operate full‑stack applications and services on Pinterest’s enterprise infrastructure to support our Sales, Marketing, and Finance teams, from backend services and APIs through integrations and user‑facing workflows. Lead the technical design and implementation of GenAI/ML‑powered services and pipelines that automate and augment enterprise workflows (for example, summarizing sales interactions, enriching account data, or surfacing intelligent recommendations), including clear evaluation frameworks, observability, and validation guardrails. Own the end‑to‑end software development lifecycle for th

JavaScriptPythonJavaNode.js
B
📍 Berkeley, United States
✓ Quality checkedCompany trend +515.8%

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

AWSAzureDockerKubernetes
V
📍 United States· Full-time
✓ Quality checkedCompany trend -88.6%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Core Platform team provides the foundational infrastructure that powers all engineering at Vanta. We're expanding upmarket to support enterprise customers, which requires strategic investment in platform systems that ensure security, reliability, and developer productivity at scale. As we expand upmarket to support enterprise and regulated customers, we’re investing heavily in platform capabilities that scale securely while reducing cognitive load for product teams. As the Engineering Manager, Core Platform at Vanta, you'll own the foundational infrastructure that every engineer builds on, ensuring it scales with company growth while remaining fast, simple, and reliable. This team’s ownership spans shared services infrastructure, observability and monitoring, datastore management, and async work systems. Our Engineering Managers develop and grow high-performing teams that deliver significant value to our customers and enable our business to scale. This role sits at the intersection of technical architecture and team development, with real authority to set direction and grow a world-class platform team. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an Engineering Manager at Vanta: Lead and grow high-performing platform engineerin

MongoDBAWSRestAI

Related career options

Similar roles with stronger pay

Client Service Associate

Demand 46/100 · 8 jobs

$840K – $840K/yr

Salary →

$840K – $840K/yr

Salary →
Director of Product

Demand 43/100 · 6 jobs

$382.5K – $382.5K/yr

Salary →
Physical Design Engineer

Demand 43/100 · 8 jobs

$300K – $300K/yr

Salary →
Sr. Engineer

Demand 42/100 · 7 jobs

$300K – $300K/yr

Salary →
Senior Director

Demand 43/100 · 22 jobs

$278.9K – $278.9K/yr

Salary →
🔔

Get new lead software engineer infrastructure jobs in United States by email

Daily job updates · Unsubscribe anytime