Jobiba hiring network

Network Reliability Engineer Jobs

1,954 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current network reliability engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

Location Details: At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operat

pythonawsci/cd
View job →

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team. As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity. What excites you Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling. Architect and manage the SLO and error-budget

awskubernetesai
View job →

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h

pythonkuberneteslinux
View job →
DC
Diligent Corporation
📍 New York• Full-time• From $131K/yr
15 days ago

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

pythonawsazure
View job →

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What You’ll Do Work with various teams to design and implement scalable, and reliable network solutions Maintain a highly available cloud infrastructure edge for the Okta identity platform C

pythonawsdocker
View job →
G
Godaddy
📍 Bulgaria• Full-time
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. As a Senior Site Reliability Engineer, you'll take direct ownership of production services — from initial design through day-to-day operation — while partnering with product, engineering, and security teams to build and maintain business-critical systems. In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs. You will also gain hands-on experience with intelligent tooling in real-world workflows. What you'll get to do... Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana Lead incident response end-to-end — performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation — validating all outputs before use Your experien

pythondockerkubernetes
View job →
O
Okta
📍 Washington• Full-time• From $165K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Technology, Data and Intelligence Team Message Okta’s Technology, Data and Intelligence (TDI) team delivers the systems, tools, and services that power internal operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale. What you'll be doing Secure Cloud Infrastructure & Pipelines: Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments. Cross-Functional Collaboration & Advocacy: Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions). Technical Documentation & Operations: Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to en

pythonawskubernetes
View job →
O
Okta
📍 Bellevue, Washington; Chicago, Illinois; New York, New York; San Francisco, California; Washington, DC• Full-time• From $194K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Team The Site Reliability team is dedicated to architecting and owning the foundational infrastructure tooling and CI/CD platforms that support Okta’s SRE ecosystem. In this development-focused role, you will leverage a modern tech-stack to build durable, automated systems that maximize platform reliability and engineering velocity. The ideal candidate is someone who enjoys analyzing systems and identifying areas of opportunity to improve system performance, availability and capacity. They are part systems administrator, part network administrator, and part developer. What you’ll be doing Maintain a highly available cloud infrastructure edge for the Okta identity platform Automate AWS infrastructure with Terraform and/or Chef Evolve the system by introducing changes to improve efficiency, scalability, and velocity What you’ll bring to the role 8+ years of operations experience configuring, deploying, monitoring and troubleshooting applications and

pythonawsdocker
View job →
O
1mo ago

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

pythonawslinux
View job →
F
15 days ago

Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment. Our customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast-growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security. Backed by world-class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people-centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI-ready operations. Forward is currently seeking experienced Java developers to work as part of our Network team. Responsibilities Help bring the best ideas from the software development world into the networking industry. Contribute to our code base, systems and software architecture as a member of our engineering team. Help create and optimize network device models for different device vendors and protocols. Help create infrastructure needed to configure, collect and test network devices. Work with peers who are experts in Networking, Distributed Systems, Big Data and Search. Requirements 5+ years of work experience in software development 3+ years of work experience with Java BS in Computer Science or related degree Solid software engineering experience with large code bases Basic understanding of networking and TCP/IP. Strong verbal and written communication skills. Nice to haves Working knowledge of how switches, routers, firewalls or load balancers work. Experience working with networking protocols such as BGP/OSPF/IS-IS, IPv4/IPv6, MPLS, VLAN, VXLAN, etc. This position is a re

javagitai
View job →

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role We’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, primarily across on-prem environments for the US Government. Forward Deployed Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll travel to various locations where you will be the expert for Palantir’s infrastructure, helping partner teams build & configure their hardware and network for software to operate reliably within. We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

pythonawsazure
View job →
TI
TextNow, Inc.
📍 San Francisco• Full-time• $136.3K – $273.9K/yr
15 days ago

TextNow is on a mission to make communications affordable and accessible for everyone. As a full MVNO operating our own mobile core network over LTE and 5G NSA, we have the unique advantage of controlling our network infrastructure end-to-end. We operate the HSS, PGW, and other critical network functions, giving us the flexibility to innovate and deliver exceptional service to millions of users. About the Role Join us in our mission to break down barriers to communication and free the flow of conversation for people everywhere. T extNow is looking for a new SecOps team member to secure, monitor , and enable automated response within our infrastructure. What You’ll Do Ensure Secure & Reliable Systems: Design, implement, and maintain security-focused infrastructure to protect TextNow’s services while ensuring reliability and scalability. Security Automation & Infrastructure as Code: Develop and enforce best practices using Terraform, Ansible, Crowdstrike , and AWS security tools , ensuring secure configurations, automated compliance checks, and infrastructure as code. Threat Detection & Incident Response: Participate in an on-call rotation to respond to security incidents, investigate vulnerabilities, and implement proactive measures to prevent future threats. Work closely with engineering teams to remediate security risks. Monitoring & Logging for Security: Improve observability by implementing security monitoring solutions, logging best practices, and alerting mechanisms to detect anomalies and suspicious activity. Access Control & Identity Management: Manage IAM roles, permissions, and policies to ensure least privilege access and enforce security controls across cloud and internal systems. Collaboration & Security Advocacy: Wo

awsci/cdgit
View job →

Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. mthree has an exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user input

pythonsqlmysql
View job →

Want to work in technology at an investment bank? Paid graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. mthree has an exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alumni will have access to the Columbia Engineering network and ongoing training. What you'll do: Production support plays a vital role in enterprise technology, from algorithmic trading engines to regulatory reporting. Think of it as healthcare for technology. As a production support analyst with mthree, you’ll be on a shared mission to look after the technical systems and processes other teams rely on. How the Alumni program works: Apply via this job advert. Complete our assessment process. Get trained at mthree Academy in an online class for 4-8 weeks with other graduates. Join a mthree client for 12-24 months while receiving support and salary increases every 12 months. The vast majority then convert to permanent employees with the client at the end of the program. What you’ll learn at the mthree Academy: How to discuss production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading, derivatives, currency, treasury, regulation, and risk. How to write a basic computer program in Python, including user

pythonsqlmysql
View job →
🔔

Get new network reliability engineer jobs by email

Daily job updates · Unsubscribe anytime