Jobs in United States

Lead Network Reliability Engineer in United States

2,434 active opportunities · Updated October 2026

Explore current lead network reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

B
📍 Berkeley, United States
✓ Quality checkedCompany trend +515.8%

Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul

AzureTerraformAnsibleSap
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We lead complex technical programs that help Plaid scale its engineering platform. We partner across Engineering, Infrastructure, Data, Security, ML, Legal, and Product to deliver company-wide technical initiatives that improve reliability, scalability, and developer productivity. You'll lead strategic technical programs from planning through execution. You'll partner with engineering leaders to align stakeholders, manage dependencies, drive decisions, and ensure successful delivery of complex initiatives. You'll work across a variety of technical domains, adapting quickly to new challenges and helping teams execute effectively. As a Technical Program Manager, you will lead high-impact, cross-functional initiatives. As a generalist, you may work on a variety of programs. An example is one that strengthens Plaid's data and machine learning platforms. You will partner with engineering, product, data, legal, privacy, and business stakeholders to drive complex technical programs from planning through execution. Your work will help improve data governance, modernize machine learning infrastructure, and accelerate the adoption of trusted, high-quality datasets that power analytics, artificial intelligence

AWSMachine LearningAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

AWSRestAIRust
L
📍 Lynnwood, WA, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos is seeking a Test and Integration Engineer to lead cutting-edge testing and validation efforts within the Undersea Systems Division (USD) . This role is a unique opportunity to drive innovation in testing underwater vehicle systems, maritime sensors, subsea telemetry, and ISR solutions that support critical defense and national security missions. You'll contribute to a multidisciplinary team focused on testing, integrating, and validating advanced maritime technologies, ensuring system reliability and performance from prototype design to full-scale system deployment for ongoing Navy missions. Leidos’ Undersea Systems Division is a recognized leader in C4ISR technologies, delivering innovative, mission-critical solutions across sensor networks, unmanned systems, and tactical platforms . We’re known for achieving “industry firsts” in the most challenging maritime domains. Join us and be part of a world-class team delivering unmatched solutions for today's most pressing maritime missions. Why Join Us? Make an Impact : Your work will directly support U.S. maritime dominance and national security. Lead Innovation : Be at the forefront of applying innovative technology and autonomy to real-world maritime systems. Work with Experts : Collaborate with a top-tier team of engineers, scientists, and technicians located across the U.S. Shape the Future : Influence both the strategic and tactical direction of next-generation subsea technologies. What You’ll Do Test and Validate: Write, develop and execute comprehensive test plans, procedures, and protocols to ensure system functionality, reliability, and compliance with requirements. Integrate systems and conduct hands-on testing: Write, develop, and execute integration plans to bring complex systems together. Perform field</b

M
📍 O Fallon, Missouri, United States
✓ High-confidence listingCompany trend +212.5%
Quick readStrong listing-quality and freshness signals

Our Purpose Mastercard powers economies and empowers people in 200&#43; countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Job Overview: Responsible for the analysis, design, development, testing, and delivery of secure, scalable software solutions. Define requirements for new applications and customization adhering to Mastercard standards, processes, and best practices. Develop, customize, and test applications to integrate to Mastercard specifications. Provide leadership, mentoring, and technical training to other team members. Major Accountabilities • Plan, design, architect, and develop secure, scalable, and maintainable technical solutions and alternatives to meet business requirements in adherence with Mastercard standards, processes, and best practices • Lead day-to-day system development and maintenance activities of the team to meet service level agreements (SLAs) and create solutions with a high level of innovation, cost effectiveness, quality, reliability, and faster time to market. • Accountable for the full systems development life cycle including creating high-quality requirements documents, use cases, designs, and other technical artifacts including but not limited to detailed test strategies, performance benchmarking, release rollout and deployment plans, contingency/back-out plans, feasibility studies, cost and time analysis, and detailed estimates. • Design, develop, test, dep

JavaDockerGitAI
M
📍 O Fallon, Missouri, United States
✓ Quality checkedCompany trend +212.5%

Our Purpose Mastercard powers economies and empowers people in 200&#43; countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team responsible for building the authentication and security solutions that help protect digital interactions across Mastercard's Identity Solutions platform. As a Senior Software Engineer, you will design, develop, and support secure, scalable, and high-performing applications that enable trusted identity verification, authentication, and fraud prevention capabilities. In this role, you will take ownership of complex technical challenges, contribute to software design and architecture decisions, and partner closely with product, security, and platform teams to deliver reliable, production-ready solutions. You'll play a key role in advancing engineering excellence through secure development practices, system reliability, automation, and continuous improvement while mentoring other engineers and influencing technical direction across the team. Role •Design, build, test, deploy, and maintain scalable, cloud-native applications and microservices •Develop REST APIs using Java and Spring Boot, focusing on performance, scalability, and reliability •Translate requirements into well-structured designs and architecture, ensuring maintainability and security •Lead and contribute to system design discussions, aligning with architectural standards and best practices<

JavaAWSAzureGCP
O
📍 United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets

AWSAzureKubernetesCI/CD
O
📍 United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches

AWSAzureGCPKubernetes
O
📍 United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI is evaluating multiple infrastructure pathways, including powered land, colo/BTS, and NeoCloud opportunities. The Site Readiness & Development team provides the diligence layer needed to compare opportunities, identify risk, and support credible deployment decisions across those pathways. About the Role The NeoCloud & Colo Due Diligence Lead will evaluate third-party infrastructure opportunities where OpenAI is considering deployment through NeoCloud, colo, or BTS structures. This role will focus on facility and deployment readiness, including MEP readiness, rack strategy, developer capability, facility design, power deliverability, schedule credibility, and operating assumptions. Unlike the land diligence team, this role is centered on technical and operational readiness of third-party infrastructure rather than greenfield site master planning, civil development, and entitlement strategy. This is an individual contributor lead role and does not have direct reports initially. The role determines whether each opportunity is fit-for-use and fit-for-service against OpenAI facility, rack, power, network, reliability, and operational standards; identifies material deficiencies and tracks remediation with developers/operators; and evaluates commissioning, validation, AHJ/code, and deployment interfaces such as structured cabling, network readiness, and high-density rack support where relevant. Key Responsibilities Lead diligence on NeoCloud, colo, and BTS opportunities across technical and operational readiness dimensions. Assess each opportunity against OpenAI facility, rack, power, network, reliability, and operational standards to determine deployment fit. Validate MEP readiness, rack deployment strategy, facility design assumptions, power deliverability, and schedule credibility. Identify material deficiencies and work with developers/operators to define remediation plans, owners, timing, and residual risk. Review reliability, availabilit

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About the Team Our Network Enablement and Access team works to unlock the potential of Plaid's network by broadening and deepening our connections with data partners. We build the capabilities that help data providers participate in the network, strengthen the quality and reliability of those connections, and enable great products and experiences for Plaid's customers. Within Network Enablement and Access, the Data Supply Traffic and Health team owns how Plaid's requests flow to the data providers we depend on. We manage the load placed on each provider, the constraints that shape our access, and the fair allocation of capacity across Plaid's products and new initiatives. As paid access expands across the network, we also work to keep that traffic reliable, efficient, and cost-effective. As the Product Manager for Data Supply Health and Traffic, you will establish and lead a new product area at the foundation of every Plaid product. You will define how Plaid allocates constrained provider capacity, scales traffic across products, and manages the economics of paid data access. You will also optimize for data freshness, balancing timeliness with provider capacity and cost so Plaid's

P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s Account Verification team builds the foundation of trust for open finance. We help fintechs and financial institutions connect and verify bank accounts securely so that money can move safely and instantly. Account Verification is the entry point for all of Plaid’s payment-focused consumer experiences and is a critical building block for our customers. The team obsesses over creating seamless verification journeys that balance speed, reliability, and security, enabling consumers to confidently connect to the financial ecosystem. As a PM for Account Verification, you’ll own one of Plaid’s most critical and high-impact product areas. You’ll lead the evolution of our verification platform across Auth, Balance, and Identity Match, defining how millions of people and businesses connect their financial accounts every day. We are looking for a high-ownership builder who thrives in ambiguity, loves building with customers, and is excited to define what’s next for one of Plaid’s most established and strategically important product lines. You’ll set vision and strategy, drive execution across a cross-functional team, and shape how Plaid competes in an increasingly complex and, eventually, AI-driven ver

AWSAIGoRust
M
📍 O Fallon, Missouri, United States
✓ High-confidence listingCompany trend +212.5%
Quick readStrong listing-quality and freshness signals

Our Purpose Mastercard powers economies and empowers people in 200&#43; countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Technical Program Manager Overview As a Senior Technical Program Manager at Mastercard, you’ll bring your expertise with conceptualizing, coordinating and driving technology projects, and coaching Agile practices in high-performing, self-organizing development teams. We’re building a global B2B, microservices-based platform to help businesses of all sizes streamline how they manage payments when buying or selling products & services. As a global business, the projects you lead for Mastercard will deliver software operating at massive scale requiring a focus on performance, security, and reliability. This role will support our Network Solutions team within Payment Networks, assisting the development team with building out software solutions for Mastercard. Role: • Dive as deep as you want into the tech stack, the integration patterns, the organizational capabilities, and the company wide assets that can be leveraged to provide technical solutions to customer problems. • Contribute to the strategies, design choices, and even the cloud infrastructure necessary to build comprehensive and achievable execution plans to deliver high-profile new features and capabilities for our customers. • Drive the execution of an initiative that may span multiple teams and integrations, reporting meani

AIRecruitment
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is looking for a hands-on Solutions Architect Manager to lead a team of GPU, networking & software solution architects and engineers. Do you want to build and lead a group that designs, debugs, and deploys new AI hardware and software technologies into production in customer data centers? As part of the NVIDIA SA organization, you will drive people and technical leadership for end-to-end solutions deployments at some of NVIDIA's most strategic technology customers, while directly contributing to designs and deep-dive debugging and shaping our product roadmap with customer feedback. What you will be doing: Recruit & manage a team of solutions architects, system/network and software engineers focused on large-scale GPU and AI networking deployments. Set priorities, allocate resources, mentor, and ensure high-quality customer delivery across multiple concurrent projects - while remaining directly involved in key technical reviews, design decisions, and critical debug efforts. Provide deep subject-matter expertise in advanced GPU and network systems and serve as the senior technical point of contact for strategic customers. Personally lead and guide complex compute/network configuration and performance debugging, working side-by-side with your team to deliver performant, reliable clusters. Guide your team as they lead network / compute / software architecture discussions, and support server, network, and cluster bring-up, including on-site data center work where needed. Systematically collect and synthesize customer-specific requirements across your portfolio. Partner with GPU/Network Systems Engineering, Product Management, and Sales to influence roadmap priorities and packaging of reference designs and solutions. Demonstrate SME in advanced GPU & network systems and be a trusted technical advisor to NVIDIA's strategic customers. Bring customer-sp

G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

PythonAIGoDevOps
C-
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$300K – $350K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are looking for a visionary, customer-obsessed and technology-focused product lead with expertise in mobile app development. In this role, you will be instrumental in shaping the future of CLEAR’s mobile app experience, delivering impactful solutions that drive user satisfaction and business growth moving CLEAR’s mobile app from a collection of features into an intelligent, trusted travel companion that helps members prepare, navigate, and recover throughout their journey. You will build a mobile App experience that is the most helpful digital layer of the travel day and be accountable to driving adoption. You will connect CLEAR’s identity capabilities, real-time travel context, airport operations, member relationships, and partner-enabled services into experiences that are personalized, proactive, and genuinely useful before, during, and after travel. This includes helping members understand airport conditions, find the right path through the airport, discover food and lounge options, and access relevant CLEAR and partner services in the moments that matter. What you'll do: Establish CLEAR’s Mobile app strategic vision and product roadmap in alignment with enterprise goals and user needs, converting a multi-year vision into prioritized product investments, structured sequencing, and core metrics like member retention and app engagement. Transform the app into an intelligent travel companion that guides members through preparation, real-time navigation, food ordering, lounge discovery and access, and any disruptions. Collaborat

GitRestAIGo
🔔

Get new lead network reliability engineer jobs in United States by email

Daily job updates · Unsubscribe anytime