About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and
Jobs in United States
Network Reliability Engineer in United States
662 active opportunities · Updated October 2026
Showing
15 jobs
Explore current network reliability engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Job Overview: Responsible for the analysis, design, development, testing, and delivery of secure, scalable software solutions. Define requirements for new applications and customization adhering to Mastercard standards, processes, and best practices. Develop, customize, and test applications to integrate to Mastercard specifications. Provide leadership, mentoring, and technical training to other team members. Major Accountabilities • Plan, design, architect, and develop secure, scalable, and maintainable technical solutions and alternatives to meet business requirements in adherence with Mastercard standards, processes, and best practices • Lead day-to-day system development and maintenance activities of the team to meet service level agreements (SLAs) and create solutions with a high level of innovation, cost effectiveness, quality, reliability, and faster time to market. • Accountable for the full systems development life cycle including creating high-quality requirements documents, use cases, designs, and other technical artifacts including but not limited to detailed test strategies, performance benchmarking, release rollout and deployment plans, contingency/back-out plans, feasibility studies, cost and time analysis, and detailed estimates. • Design, develop, test, dep
Leidos is seeking a Test and Integration Engineer to lead cutting-edge testing and validation efforts within the Undersea Systems Division (USD) . This role is a unique opportunity to drive innovation in testing underwater vehicle systems, maritime sensors, subsea telemetry, and ISR solutions that support critical defense and national security missions. You'll contribute to a multidisciplinary team focused on testing, integrating, and validating advanced maritime technologies, ensuring system reliability and performance from prototype design to full-scale system deployment for ongoing Navy missions. Leidos’ Undersea Systems Division is a recognized leader in C4ISR technologies, delivering innovative, mission-critical solutions across sensor networks, unmanned systems, and tactical platforms . We’re known for achieving “industry firsts” in the most challenging maritime domains. Join us and be part of a world-class team delivering unmatched solutions for today's most pressing maritime missions. Why Join Us? Make an Impact : Your work will directly support U.S. maritime dominance and national security. Lead Innovation : Be at the forefront of applying innovative technology and autonomy to real-world maritime systems. Work with Experts : Collaborate with a top-tier team of engineers, scientists, and technicians located across the U.S. Shape the Future : Influence both the strategic and tactical direction of next-generation subsea technologies. What You’ll Do Test and Validate: Write, develop and execute comprehensive test plans, procedures, and protocols to ensure system functionality, reliability, and compliance with requirements. Integrate systems and conduct hands-on testing: Write, develop, and execute integration plans to bring complex systems together. Perform field</b
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview Join a team responsible for building the authentication and security solutions that help protect digital interactions across Mastercard's Identity Solutions platform. As a Senior Software Engineer, you will design, develop, and support secure, scalable, and high-performing applications that enable trusted identity verification, authentication, and fraud prevention capabilities. In this role, you will take ownership of complex technical challenges, contribute to software design and architecture decisions, and partner closely with product, security, and platform teams to deliver reliable, production-ready solutions. You'll play a key role in advancing engineering excellence through secure development practices, system reliability, automation, and continuous improvement while mentoring other engineers and influencing technical direction across the team. Role •Design, build, test, deploy, and maintain scalable, cloud-native applications and microservices •Develop REST APIs using Java and Spring Boot, focusing on performance, scalability, and reliability •Translate requirements into well-structured designs and architecture, ensuring maintainability and security •Lead and contribute to system design discussions, aligning with architectural standards and best practices<
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
JLL empowers you to shape a brighter way . Our people at JLL are shaping the future of real estate for a better world by combining world class services, advisory and technology for our clients. We are committed to hiring the best, most talented people and empowering them to thrive, grow meaningful careers and to find a place where they belong. Whether you’ve got deep experience in commercial real estate, skilled trades or technology, or you’re looking to apply your relevant experience to a new industry, join our team as we help shape a brighter way forward. Automation Engineer – JLL What this job involves: We are seeking an experienced Automation Engineer to design, develop, and implement automation control systems for industrial processes and warehouse distribution equipment. The role requires strong knowledge of engineering principles, programming, and control system technologies, with a focus on improving the reliability and performance of conveyors, sortation systems, scanners, cameras, print-and-apply systems, and SCADA devices. All work must follow established policies and procedures, with safety as a top priority. What your day-to-day will look like: Serve as site technical expert in automation control systems and mentor Apprentices to meet safety and technical standards. Design, develop, implement, and optimize control systems and software; maintain and troubleshoot equipment including PLC/PC controllers and industrial networks. </
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Software Engineer II Overview Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. The Fraud Products team (part of O&T) is developing new capabilities for MasterCard's Decision Management Platform, which serves as the core for multiple business solutions to combat fraud and validate cardholder identity. Our patented Java-based platform processes billions of transactions per month in tens of milliseconds using a multi-tiered, message-oriented approach for high performance and availability. MasterCard software engineering teams leverage Agile development principles, advanced development, design and test automation practices, and an obsession over security, reliability, and perfo
What we're building Mutiny is the self-improving AI infrastructure for GTM teams to execute faster and close more revenue. Our ambition is to do for revenue velocity what Cursor and Claude Code did for engineering velocity. With Mutiny, everyone in sales and marketing gets a bench of GTM athletes that handle any work across their revenue motion and learn from what's actually moved their deals. In April we re-launched the product as an agent-first platform. Anthropic showcased us as a leader in AI GTM. MRR is growing more than 70% month-over-month, with customers like Uber, Rippling, and Snowflake. We're backed by Sequoia, YC, and Insight, and we're building a generational company. The opportunity Most engineers spend their career making predictable systems faster. You'll spend yours making non-deterministic ones trustworthy. As a senior engineer on our AI product team, you'll architect the Campaign Builder and Agent experiences marketers and sellers open every day to go from idea to personalized assets in minutes. You'll partner directly with product, design, and the founders to define what an agent-first GTM platform should feel like, and your calls on architecture, evals, and guardrails compound across thousands of customer accounts. This role is in person in New York City, five days a week, and we ship weekly. What you'll own The core agent surfaces. Architect and ship the Campaign Builder and Agent experiences end-to-end. Frontend, backend, prompts, evals, the whole stack. Reliability on top of LLMs. Make non-deterministic models feel deterministic at the surface. Build the retries, fallbacks, and orchestration so the customer never sees the failure mode. Evals and guardrails. Define how we measure quality, catch regressions, and keep brand and tone consistent across thousands of customer accounts. Speed and feel. AI products live or die by latency and the loop between intent and output. You'll obsess over both, and use coding agents and agent networks to ship f
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
$135K – $225K/yr
About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are We are seeking an experienced DevOps Engineer to join our growing team and play a pivotal role in designing and building our platform and infrastructure as we continue to scale our product and user base. As a part of our team, you will be working in a dynamic, fast-paced environment to ensure the reliability, scalability, and performance of our systems, while focusing on service architecture and deployment, query optimization, distributed systems, data and machine learning infrastructure, and security and authentication. Most importantly, you are excited to be part of a mission-oriented, fast-paced, high-growth startup that can create a lasting impact. You will: Partner with product teams to architect, design, and build the foundational infrastructure for our products. Design, develop, and deploy highly available and scalable Multi-tenant SaaS solutions on any one of the public cloud networks like AWS, Azure and GCP. Leverage technologies such as Kubernetes, Helm, Terraform, and Istio to achieve infrastructure resilience. Drive the automation of infrastructure tasks, from provisioning to configuration management and deployment, utilizing tools like Terraform, Ansible, a
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About the Team Our Network Enablement and Access team works to unlock the potential of Plaid's network by broadening and deepening our connections with data partners. We build the capabilities that help data providers participate in the network, strengthen the quality and reliability of those connections, and enable great products and experiences for Plaid's customers. Within Network Enablement and Access, the Data Supply Traffic and Health team owns how Plaid's requests flow to the data providers we depend on. We manage the load placed on each provider, the constraints that shape our access, and the fair allocation of capacity across Plaid's products and new initiatives. As paid access expands across the network, we also work to keep that traffic reliable, efficient, and cost-effective. As the Product Manager for Data Supply Health and Traffic, you will establish and lead a new product area at the foundation of every Plaid product. You will define how Plaid allocates constrained provider capacity, scales traffic across products, and manages the economics of paid data access. You will also optimize for data freshness, balancing timeliness with provider capacity and cost so Plaid's
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s Account Verification team builds the foundation of trust for open finance. We help fintechs and financial institutions connect and verify bank accounts securely so that money can move safely and instantly. Account Verification is the entry point for all of Plaid’s payment-focused consumer experiences and is a critical building block for our customers. The team obsesses over creating seamless verification journeys that balance speed, reliability, and security, enabling consumers to confidently connect to the financial ecosystem. As a PM for Account Verification, you’ll own one of Plaid’s most critical and high-impact product areas. You’ll lead the evolution of our verification platform across Auth, Balance, and Identity Match, defining how millions of people and businesses connect their financial accounts every day. We are looking for a high-ownership builder who thrives in ambiguity, loves building with customers, and is excited to define what’s next for one of Plaid’s most established and strategically important product lines. You’ll set vision and strategy, drive execution across a cross-functional team, and shape how Plaid competes in an increasingly complex and, eventually, AI-driven ver
Other cities to consider
More places hiring for this role
Get new network reliability engineer jobs in United States by email
Daily job updates · Unsubscribe anytime