Jobiba hiring network

Senior Network Reliability Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior network reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

C
Cloudflare
📍 Hybrid• Full-time• Hybrid
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX About the Role You’ll help define how machine learning models run across Cloudflare’s global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You’ll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and effic

pythonawsmachine learning
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Chemical and Slurry Systems Engineer within Micron’s Global Facilities and Construction Team, you will deliver engineering expertise supporting the planning, design, construction, operation, and maintenance of facilities systems across Micron’s global manufacturing network. You will participate in capacity and scenario planning, evaluate impact to facilities infrastructure, and serve as a technical domain expert throughout all project phases. Your work ensures environmental, safety, regulatory, and code compliance while driving system reliability and operational excellence. Responsibilities Develop safe, reliable designs for Chemical and Slurry systems and ensure compatibility of materials of construction with all chemicals used. Create, review, and maintain corporate equipment specifications, engineering standards, and design guidelines. Advise sites on system capacity, load projections, and scenario‑driven impacts. Provide technical feedback on infrastructure additions or modifications to support manufacturing and technology roadmap changes. Supply Micron standards to design partners and conduct timely reviews of design packages, construction documents, and engineering work. Support project cost and schedule development and validate scope alignment with collaborator requirements. Coordinate with Facilities, Design, Construction, Procurement, and Manufacturing to ensure clear communication and project alignment. Provide technical guidance during construction,

aiprocurementrecruitment
View job →

About Inspira Education Inspira Education Group is one of the fastest-growing edtech startups in the US. We started with a simple mission to democratize access to high-quality coaching so that every student in the world has an equal opportunity to access the best opportunities. As the world’s leading network of top admissions coaches in medical, legal, business, and college studies, we’re building software and services in one place—disrupting long-entrenched application processes with products and experiences that strive to provide an equal platform for candidates from diverse backgrounds worldwide. As one of the fastest-growing edtech firms in the world, we are backed by some of the leading venture capital firms and investors in the world, including Zeev Ventures, Quiet Capital, Craft Ventures and Jeff Fluhr (Founder of Stubhub). About the role We’re looking for a strong full-stack engineer who can own the complete product development process: understand a business problem, define the solution, design the user experience, build the software, and improve it after launch. You’ll work closely with leadership and business teams, combining hands-on engineering with product management and design responsibilities. You should be highly effective with AI coding tools and have the technical depth to independently review, debug, secure, and maintain everything you ship. This is an in-person role requiring 5 day/week in our NYC office. What you’ll own Translate business needs and user feedback into product requirements, user flows, prototypes, and prioritized development plans. Design and build polished applications across the front end, back end, database, and integrations. Make architecture decisions and scope releases that balance speed, reliability, and future maintainability. Use AI tools throughout development to accelerate implementation, testing, debugging, and documentation. Own deployment, production monitoring, incident resolution, and ongoing improvemen

javascripttypescriptpython
View job →

About Inspira Education Inspira Education Group is one of the fastest-growing edtech startups in the US. We started with a simple mission to democratize access to high-quality coaching so that every student in the world has an equal opportunity to access the best opportunities. As the world’s leading network of top admissions coaches in medical, legal, business, and college studies, we’re building software and services in one place—disrupting long-entrenched application processes with products and experiences that strive to provide an equal platform for candidates from diverse backgrounds worldwide. As one of the fastest-growing edtech firms in the world, we are backed by some of the leading venture capital firms and investors in the world, including Zeev Ventures, Quiet Capital, Craft Ventures and Jeff Fluhr (Founder of Stubhub). About the role We’re looking for a strong full-stack engineer who can own the complete product development process: understand a business problem, define the solution, design the user experience, build the software, and improve it after launch. You’ll work closely with leadership and business teams, combining hands-on engineering with product management and design responsibilities. You should be highly effective with AI coding tools and have the technical depth to independently review, debug, secure, and maintain everything you ship. This is an in-person role requiring 5 day/week in our NYC office. What you’ll own Translate business needs and user feedback into product requirements, user flows, prototypes, and prioritized development plans. Design and build polished applications across the front end, back end, database, and integrations. Make architecture decisions and scope releases that balance speed, reliability, and future maintainability. Use AI tools throughout development to accelerate implementation, testing, debugging, and documentation. Own deployment, production monitoring, incident resolution, and ongoing improvemen

javascripttypescriptpython
View job →
N
Nvidia
📍 Remote, Poland• Remote
10 days ago

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, passionate, and self-motivated, we want to hear from you! We are looking for an experienced networking software engineer. An awesome candidate is highly technical who is also comfortable with dealing with enterprise customers. You will join a team of Solution Engineers focused on the Mellanox Networking, DGX Platforms, Container Orchestrators, Deep Learning containers, and other Enterprise related system software. SW Solution Engineers spend approximately 50% of their time helping customers with their most complex problems and 50% of their time doing R&D related work. This individual should have proven grasp of datacenter and networking technologies, to provide comprehensive solutions for complex installations, maintenance, or operations for a broad scope of leading-edge networking products. What you'll be doing: Take ownership and drive customer issues with Ethernet or InfiniBand network adapter/DPU deployments from inception to resolution. Develop features and tools as part of solution engineering efforts to support all Enterprise Service offerings including but not limited to Networking products. Work with NVIDIA Enterprise customers and internal users to improve the availability, reliability, and overall experience of working with NVIDIA Networking products. Bring independent analysis, communication, and problem-solving to customer experience. Collaborate with engineering to document, recreate and solve issues. What we need to see: BSc in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience). 8+ years system software developm

REMOTEkuberneteslinux
View job →
G
18 days ago

Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team The Network Security team is responsible for securing GoDaddy's global hybrid infrastructure across data centres, cloud, edge, and remote-access environments. We partner across Security, Infrastructure, Cloud, and Engineering teams to build scalable, resilient, and secure solutions that support the business. As a Principal Network Security Engineer, you'll act as a senior technical leader, helping define network security strategy, influence architecture across teams, and drive security outcomes at enterprise scale through technical expertise, systems thinking, and cross-functional leadership. What you'll get to do... Define and drive network security architecture across hybrid environments, including data centres, cloud, edge, and remote-access technologies Design trust boundaries, segmentation strategies, secure connectivity patterns, and network controls that reduce risk and improve security posture Lead complex technical initiatives, migrations, and architectural decisions while balancing security, reliability, performance, and operational requirements Establish scalable approaches for policy governance, automation, monitoring, telemetry, and security control validation Partner across engineering organizations to drive large-scale initiatives, mentor engineers, and influence technical direction through architecture reviews and technical leadership Your experience should include... 10+ years of experience in Network Security Engineering, Network Architecture, or Security Engineering, including ownership of enterprise-scale secu

awsci/cdai
View job →
P
1mo ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard

sqlmysqlaws
View job →
P
Plaid
📍 New York• Full-time
1mo ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard

sqlmysqlaws
View job →
P
Plaid
📍 San Francisco• Full-time
1mo ago

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid’s mission is to unlock financial freedom for everyone by making money movement and access to financial data simple and secure. As a Software Engineer, you will design and build the systems that power how millions of people connect to their finances. You will work across the stack, from reliable backend services and APIs to intuitive applications that bring those systems to life. You will collaborate with engineers, product managers, and designers to ship products that make financial services more accessible and transparent. At Plaid, engineers take ownership early, grow quickly, and see their work reach millions of users. Responsibilities: Design & Development: Build and maintain backend services with a focus on performance, reliability and scalability. Collaboration: Work closely with product managers and other stakeholders to define and implement new features that meet product and customer needs. Code Quality: Write clean, maintainable and efficient code. Testing & Debugging: Develop automated tests to ensure the quality and reliability of the codebase. Troubleshoot and resolve issues. Engage in hands-on coding and architectural design, setting and maintaining high technical standard

sqlmysqlaws
View job →
C
Cloudflare
📍 In Office• Full-time• $194K – $266K/yr
1mo ago

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. Available Locations: Austin, TX or San Francisco, CA About the Role AI inference is becoming core infrastructure. Every serious application will need access to many models, across many providers, with reliability, observability, security, cost control, and routing built in from the start. AI Gateway is Cloudflare’s bet that this layer should exist at the network edge: close to users, close to compute, and simple enough that a developer can adopt i

typescriptawsrest
View job →
L
Lyft
📍 San Francisco• Full-time
1mo ago

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. The Lyft Business Product Platform team builds the systems and experiences that power Lyft's B2B products — enabling companies, organizations, and their employees to seamlessly access Lyft's transportation network. We sit at the intersection of product and platform, owning both the customer-facing features and the underlying infrastructure that makes them reliable at scale. Our work directly impacts how businesses integrate with Lyft, how admins manage their programs, and how millions of riders get where they need to go. Responsibilities: Drive architecture and technical design for systems that are highly available, scalable, and built to last — not just for today's requirements but for where the product is heading Own features end-to-end: from shaping the technical spec and design through to production rollout and operational health Think critically about how AI capabilities can be incorporated into Lyft Business products to improve the experience for business admins and riders — and bring that perspective into roadmap and architecture conversations Make well-reasoned trade-off decisions and communicate them clearly to peers, leads, and cross-functional partners Write clean, well-tested, maintainable code and hold a high bar for the same in code reviews Partner across engineering, product, and design to align on direction and get buy-in on technical approaches Proactively engage in incident response, contributing both to resolution and to long-term reliability improvements Grow the team's technical culture through design reviews, tech talks, and mentorship Experience: 5+ years of software engineering experience, with a track record of designing and shipping production systems at scale Strong system design instincts — you can reason through distributed systems trade-offs, identify failure modes, and

sqlawsazure
View job →

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are You are an experienced Infrastructure Engineer Engineer who owns backend infrastructure end to end. You design multi-tenant, microservices-based systems that other engineering teams build on, and you make deliberate architectural tradeoffs around consistency, latency, scale, and cost. You are comfortable going deep — service mesh internals, database internals, distributed-systems failure modes — and equally comfortable defining the reliability and security contracts an enterprise AI platform depends on. Responsibilities Design, own, and evolve scalable microservices architectures on Kubernetes across GCP, Azure, and AWS, including multi-tenant isolation (namespaces, network policies, per-tenant resource quotas and RBAC). Build core platform and data-plane components in Golang and Python — data ingestion, knowledge-base indexing and vector/graph search, application connectivity, workflow automation, and ML operations — against explicit latency and throughput SLOs. Own service-to-service communication: gRPC/protobuf API contracts, service mesh (Istio/Linkerd), load balancing, retries, timeouts, and circuit breaking. Make and document architectural tradeoffs — partitioning

pythonsqlaws
View job →

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl

pythonlinuxai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →
SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
14 days ago

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

awsazuregcp
View job →
🔔

Get new senior network reliability engineer jobs by email

Daily job updates · Unsubscribe anytime