Jobs in United States

Lead Network Reliability Engineer in United States

2,434 active opportunities · Updated October 2026

Explore current lead network reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

PythonAWSAzureCI/CD
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team. As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity. What excites you Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling. Architect and manage the SLO and error-budget

AWSKubernetesAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. Our Stargate program develops and deploys massive, state-of-the-art data center campuses in partnership with industry leaders today—and through future OpenAI infrastructure projects tomorrow. We design for scale, speed, and reliability, and we need experienced technicians who can translate network blueprints into physical reality. About the Role: We are seeking a Senior Data Center Networking Technician who thrives in fast-moving build environments and is eager to roll up their sleeves during active datacenter deployments. Your first assignment will focus on the physical bring-up of network infrastructure at a large partner-operated campus, collaborating with partner teams and their delivery vendors to achieve agreed performance and reliability targets. As that campus reaches steady state, you will transition to lead network deployment for future OpenAI data center projects, defining standards and guiding implementation across multiple locations. Candidates must be able to sit onsite in Abilene, Texas 5 days per week Key Responsibilities Serve as OpenAI’s technical lead technician during the current campus build, partnering with internal engineers and external contractors on design reviews, installation plans, and acceptance criteria. Spend significant time on the data-center floor performing inspections, assisting with cable routing/termination when needed, conducting fiber testing (OTDR, power levels, continuity), and resolving installation challenges in real time. Troubleshoot and optimize cabling routes, patching, and equipment turn-up to ensure clean, reliable handoff to network operations. Contribute to design discussions and peer reviews for structured cabling and physical network layouts, providing practical field feedback to engineering teams. Develop repeatable engineering standards, as-built do

PythonAWSLinuxRest
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $152K/yr

Quick readStrong listing-quality and freshness signals

The Manager of Networking at Datadog leads the global management of network services across all Datadog offices worldwide. This role is responsible for ensuring seamless, high-performance Wi-Fi and direct internet access in our global offices and conference room technology, supporting a rapidly growing global enterprise. As Datadog continues to grow rapidly, this role plays a critical part in scaling both the team and network infrastructure to meet increasing demand. This is a hybrid role that sits in global headquarters in New York city and requires three days in the office each week with occasional travel to our offices around the world. What You’ll Do: Lead the global network engineering teams, managing both full-time employees and third-party vendors to ensure consistent and high-quality service delivery for Datadog offices worldwide. Oversee the design, deployment and scaling of office network infrastructure, including Wi-Fi (Cisco Meraki) and edge networking devices (Cisco, Palo Alto, Juniper, and others), ensuring these services operate according to Datadog's defined service level objectives. Define and implement standards, policies, and processes for network infrastructure to ensure security, reliability and scalability. Collaborate with IT Security, Enterprise Technology, and Workplace teams to align network services with broader IT and business objectives. Develop and track operational metrics for service availability, network performance, driving continuous improvement and optimization. Who You Are: An experienced people manager with at least 5+ years of leadership experience managing teams of network engineers. Proven expertise in Wi-Fi network engineering, including deep knowledge of Cisco Meraki and edge networking solutions from vendors like Cisco, Palo Alto and Juniper. Experience managing large-scale office technology projects in global enterprises with more than 10 offices and 7,000+ employees, ensuring infrastructure keeps pace with ra

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the role: We are seeking an experienced Optical Network Engineer to lead Laser related work within our optical interconnect efforts for large-scale compute systems. The role also requires broad, hands-on optical validation experience across IM/DD-based interconnects, working from lab characterization through production readiness and scaled deployment. In this role you will: Drive laser-focused requirements and technical direction within the broader optical interconnect roadmap. Lead evaluation and validation of optical components and subsystems, including laser-based elements, in lab and production-representative environments. Support end-to-end optical testing for IM/DD interconnects (e.g., module/system bring-up, characterization, debug, and readiness for scale). Work with external partners to align on development milestones, performance targets, and quality expectations. Own technical issue triage and resolution across performance, reliability, and manufacturability topics. Collaborate across internal teams to support integration, rollout, and operational success at scale. You might thrive in this role if you have: Strong experience in laser-focused optical engineering (development, validation, manufacturing readiness, or field support). Broad hands-on background with IM/DD optical technologies and optical test/debug workflows. Experience working with external suppliers/manufacturing partners and production-oriented execution. Demonstrated ability to debug complex t

AWSRestAIRust
I
📍 Arizona, Phoenix, United States
✓ Quality checkedCompany trend +315.4%

Job Details: Job Description: As one of the world's largest semiconductor manufacturers, Intel is committed to advancing every aspect of semiconductor technology, from process development and manufacturing to advanced packaging and reliability characterization. Employees within Intel Foundry are part of a global network spanning technology development, manufacturing, assembly, test, and quality organizations across both front-end silicon and advanced packaging facilities. You will join the Foundry Lab Network (FLN), a key organization within Foundry Quality, Reliability, and Labs (FQRL), located at Intel's growing Chandler, Arizona site. This site serves as the technology development hub for Intel's most advanced packaging technologies, including Hybrid Bond Interconnect (HBI), Embedded Multi-die Interconnect Bridge (EMIB-T), glass substrates, Foveros, and future advanced packaging innovations. FLN is actively expanding its laboratory capabilities and capacity to support Intel Foundry's roadmap, enabling faster technology qualification and accelerating development cycles through Quick Turn Monitoring (QTM) solutions that significantly reduce time-to-data. As the Stress Test and Reliability (STaR) Back-End Laboratory Manager, you will play a critical leadership role within FLN. You will partner closely with Packaging Technology Development, Quality and Reliability, and Failure Analysis teams to establish and enhance laboratory capabilities that support qualification and certification of next-generation packaging technologies and products. This position offers a unique combination of people leadership, technical problem solving, and organizational strategy. You will lead a team of 7-10 talented engineers, spearhead cross-functional efforts to resolve complex reliability and technology challenges, drive strong quality and operati

SQLAuditingRecruitment
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$275K – $350K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng

PythonJavaAWSKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa

AWSRestAIGo
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. Today, CLEAR is well-known as a leader in digital and biometric identification, reducing friction for our members wherever an ID check is needed. We’re looking for a Senior Software Engineer to establish our Observability framework and foundations. You will join us to accelerate building and scaling our innovative systems that support our growing identity platform. You will drive on Observability best practices to find and fix gaps in our observability and our overall systems. You will also lead practices such as load testing, capacity planning, game days, chaos testing, and incident post-mortems. What You Will Do: Embed within the Engineering pillar to deeply understand the product and implement observability across all key flows Facilitate and build load testing cases, ensuring we understand the limits and scaling factors of our services and systems Contribute to observability and support the design of new services and systems, ensuring highly reliable and scalable concepts are implemented Build and lead practices such as game days, chaos engineering, and failure analysis Build long-term capacity plans, with an eye toward reliability and cost-efficiency Who You Are: 6+ experience writing production-grade software in a modern language, such as Java and Python. Strong knowledge of distributed systems concepts (think CAP theorem), microservices architecture, and distributed tracing . Experience with modern observability systems such as Datadog. Experience with performance debugging tools and patterns. You should be able to read a f

PythonJavaGitRest
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$435K – $535K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a collaborative, strategic, and execution-focused engineering leader to shape the digital experiences that millions of people have with CLEAR. As a Director of Engineering, you will lead multiple engineering teams responsible for building and scaling our consumer products across the CLEAR mobile app, website, digital enrollment experiences, marketing technology, and Concierge platform. You'll partner closely with Product, Design, Marketing, and business leaders to create intuitive, high-performing experiences that drive acquisition, engagement, conversion, and member satisfaction. The ideal candidate combines strong technical leadership with a deep understanding of consumer product development, building high-performing teams that move quickly while maintaining quality, reliability, and operational excellence. What You'll Do: Lead and grow multiple engineering teams responsible for CLEAR's consumer experiences across mobile, web, marketing technology, and Concierge products. Define and execute the technical strategy for customer-facing applications, ensuring scalable, reliable, and performant experiences that delight millions of members. Partner closely with Product, Design, Marketing, and Business stakeholders to prioritize roadmaps that improve acquisition, conversion, engagement, retention, and member satisfaction. Drive engineering excellence by establishing best practices around architecture, delivery, quality, observability, and operational performance. Build highly collaborative relationships across Engineering, Pro

GitRestAIGo
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $244K/yr

Quick readStrong listing-quality and freshness signals

Datadog’s Cloud Networks team designs, builds, and maintains the production network infrastructure that powers everything built on top of our platform across AWS, GCP, Azure, and beyond. In this role, you’ll set technical direction for how we scale our multi-region, multi-cloud network footprint while keeping reliability and performance high. You’ll partner closely with internal teams and Cloud Service Providers to troubleshoot complex connectivity issues, integrate new networking capabilities, and improve the foundations our engineers and customers rely on. This is a high-impact opportunity to drive meaningful improvements in scale, resiliency, and cost efficiency. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design, build, and operate cloud network infrastructure across AWS, GCP, Azure, and Neoclouds in a multi-region environment. Own connectivity between clouds, customers, and developers—ensuring scalable, secure, and reliable network paths. Set clear technical direction for expanding data centers and evolving the network while maintaining stability and performance. Improve cross-site and cross-region connectivity patterns to support Datadog’s growing platform needs. Lead deep investigations into latency, packet loss, and connectivity failures – from pcap and path analysis through to escalations with cloud providers that may originate from customer support Identify and deliver network-related efficiency and cost-saving opportunities that positively impact business health. Who You Are: You have deep networking expertise. You understand BGP, route policies, path selection, prefix advertisement, and what breaks in large-scale networking. You have substantial experience designing, building, and evolving large-scale Software-Defined Networks—inclu

AWSAzureGCPAI
C-
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$275K – $350K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. The Engineering Manager, CLEAR1 Strategic Partnerships Platform (StratP) is responsible for leading the engineering team that builds partner integrations, healthcare-related workflows, identity and authentication capabilities, and reusable platform improvements that strengthen CLEAR1’s customer offerings. In this role, you’ll drive execution, technical direction, and team development across a high-leverage platform area, turning strategic partner and market needs into scalable, reliable capabilities that improve delivery quality and create long-term business value. What you’ll do: Lead, coach, and develop engineers through hiring, feedback, performance management, and career growth Drive roadmap execution by setting priorities, running strong planning cadences, and delivering against CLEAR1 business commitments Partner across Product, Design, Operations, and Engineering to translate partner and customer needs into scalable platform capabilities Lead delivery across work spanning partner integrations, healthcare-related workflows, identity and authentication capabilities, and reusable platform improvements for CLEAR1 Guide architecture, technical tradeoffs, and dependency management to improve reliability, speed, and long-term platform leverage How you’ll measure success: Improved roadmap predictability and delivery against planned commitments Clearer prioritization and stronger execution across competing demands and shared dependencies Better ownership clarity across partner-facing integrations, identity flows, and customer-facing platf

GitRestAIGo
🔔

Get new lead network reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime