The eBPF APM team is building a zero-instrumentation observability solution that automatically discovers services on every host, supports both plaintext and TLS-encrypted traffic, classifies Layer 7 protocols, decodes service-level traffic, and reports RED (requests, errors, duration) metrics. Leveraging deep expertise in eBPF, the team operates across a wide range of Linux kernel versions, distributions, and complex customer environments. In addition to low-level networking, the team solves challenges related to protocol versioning, TLS detection across diverse languages and runtimes, and resilient performance in production systems We’re looking for a senior engineer with strong systems-level thinking and a good understanding of Linux. You should be comfortable working close to the kernel, ideally with experience in eBPF, or with a strong desire to dive into it. Proficiency in C/C++/ Go is essential, and familiarity with networking protocols, TLS internals, or distributed tracing is a strong advantage. You’ll join a high-impact team tackling ambitious technical challenges—like decoding traffic across multiple protocols, and ensuring high-fidelity metrics in complex, real-world environments. You’ll be expected to lead design and implementation efforts, contribute to roadmap planning, and collaborate across teams to ensure our solution remains robust, scalable, and frictionless for our users. This role is a great fit for engineers who thrive on low-level, performance-sensitive problems, and want to shape the future of observability through cutting-edge kernel technology. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design and build core components of our zero-instrumentation APM product using eBPF and Go Develop systems to aut
Jobiba hiring network
Networking Manager Jobs
555 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current networking manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Datadog’s Cloud Networks team designs, builds, and maintains the production network infrastructure that powers everything built on top of our platform across AWS, GCP, Azure, and beyond. In this role, you’ll set technical direction for how we scale our multi-region, multi-cloud network footprint while keeping reliability and performance high. You’ll partner closely with internal teams and Cloud Service Providers to troubleshoot complex connectivity issues, integrate new networking capabilities, and improve the foundations our engineers and customers rely on. This is a high-impact opportunity to drive meaningful improvements in scale, resiliency, and cost efficiency. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Design, build, and operate cloud network infrastructure across AWS, GCP, Azure, and Neoclouds in a multi-region environment. Own connectivity between clouds, customers, and developers—ensuring scalable, secure, and reliable network paths. Set clear technical direction for expanding data centers and evolving the network while maintaining stability and performance. Improve cross-site and cross-region connectivity patterns to support Datadog’s growing platform needs. Lead deep investigations into latency, packet loss, and connectivity failures – from pcap and path analysis through to escalations with cloud providers that may originate from customer support Identify and deliver network-related efficiency and cost-saving opportunities that positively impact business health. Who You Are: You have deep networking expertise. You understand BGP, route policies, path selection, prefix advertisement, and what breaks in large-scale networking. You have substantial experience designing, building, and evolving large-scale Software-Defined Networks—inclu
We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Own the reliability architecture of the platform across regions and cloud providers Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices Set operational standards for the team: on-call quality, incident response, SLO discipline Mentor and technically develop the SRE team Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Qualifications 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms Proficiency in Python, Go, or a similar programming language Understand workload isolati
We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing Identify and configure key metrics to detect incidents and quantify service health, availability, and performance Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Mentor early-career SREs and contribute to the team’s operational practices as it grows Qualifications Strong background in software development and operating distributed systems 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues Expertise in cloud infrastructure platforms, in
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe Infrastructure is responsible for the reliability, scale, performance, and cost of Stripe's systems and the productivity and sentiment of Stripe's people. You may work on a wide variety of critical business areas including: Core Infrastructure—We're the home for Stripe's critical tier0 infrastructure systems (Compute, Networking, DocumentDB, Distributed Caching and High assurance engineering). We build the foundational platform for Stripe products and services to allow them to operate at scale. We drive reliability, availability, efficiency, and scalability of these systems. Developer Infrastructure—We're responsible for the productivity of all developers at Stripe. Ensure Stripe's engineers have a reliable, fast, and easy-to-use inner dev loop to maximize productivity while building everything from low-latency microservices to large-scale data pipelines and machine learning models. Data Infrastructure—We're responsible for offering data serving infrastructure spanning across data warehouse analytics, streaming analytics, and search capabilities. The stack is supported by a collection of internally developed large-scale distributed services and several popular open-source technologies like Trino/Presto, Apache Pinot, Hive Metastore, ElasticSearch etc. The systems we own support all of the data serving needs of high-scale services and thousands of individual Stripes across the company. Admin Platform—We empower Stripes to quickly
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role We're looking for an Optical Interconnect System Engineer to design, qualify, and deploy scalable optical connectivity for large-scale AI infrastructure. This role spans fiber-system architecture, optical-mechanical integration, validation, reliability, deployment, and serviceability. You will work with optical, mechanical, electrical, networking, manufacturing, reliability, and data-center teams to translate system needs into practical interconnect solutions. This is a hands-on role for someone who can connect design decisions with installation, qualification, troubleshooting, and long-term operational performance. In this role, you will: Define optical interconnect architectures and requirements across hardware platforms and rack-level systems. Design high-density fiber systems for performance, density, reliability, installation, and serviceability. Lead optical-mechanical integration and cross-functional design reviews. Develop test and qualification plans for optical components, modules, switching platforms, and integrated systems. Own optical loss budgets, routing guidelines, handling requirements, and serviceability criteria. Support system bring-up, deployment, troubleshooting, failure analysis, and reliability improvement. Create reusable design guidelines, interface requirements, and qualification methods. You might thrive in this role if you have: Core experience Experience desi
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for an experienced systems software engineer to help define and build the host software stack for our custom next-generation AI systems. You will work close to the hardware on performance-critical software, including Linux kernel drivers, high-throughput I/O paths, and system-scale networking and RDMA. This role spans architecture, implementation, platform bring-up, debugging, and performance optimization. You will work across hardware and software boundaries to make new systems usable end to end, from low-level device interfaces through userspace tooling and production validation. In this role you will: Design, implement, and debug host-side systems software for AI infrastructure, including Linux kernel drivers and supporting userspace components. Build and optimize software paths for high-throughput, low-latency communication, including RDMA and related networking functionality. Develop software around PCIe, DMA, NICs, accelerators, memory movement, and device interaction. Bring up new hardware platforms and diagnose complex issues across kernel, firmware, networking, and hardware boundaries. Build tooling for integration, testing, diagnostics, observability, qualification, and performance characterization. Collaborate with hardware, networking, and platform teams to define interfaces and integrate new capabilities. Work with external vendors where needed to integrate technologies and drive issues to resolution. Contribute across the systems sof
About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of AI at unprecedented scale. Through a combination of strategic partnerships and self-built campuses, we are developing and operating large-scale data center infrastructure across power, cooling, networking, compute, construction, and site operations. The scale and complexity of this infrastructure introduces a broad range of environmental, health, and safety considerations across site development, design, construction, equipment deployment, commissioning, and ongoing operations. EHS is a critical part of how we build infrastructure that is safe, resilient, compliant, and capable of operating at scale. About the Role We are seeking an EHS Lead to establish and drive environmental, health, and safety strategy across OpenAI’s rapidly expanding compute infrastructure portfolio. This role will develop the EHS framework for large-scale data center development and operations, partnering closely with engineering, construction, infrastructure delivery, facilities, operations, security, legal, environmental, and external development partners. The EHS Lead will help ensure that safety and environmental considerations are embedded into projects from early design and site development through construction, commissioning, and operations. The role will establish standards and operating mechanisms, assess and mitigate risks, oversee EHS performance across internal teams and third-party partners, and provide technical leadership on complex or high-consequence safety issues. Success requires the ability to operate strategically while maintaining strong technical depth and executional rigor in fast-moving, highly complex infrastructure environments. In this role, you will: Develop and own EHS strategy, standards, programs, and operating mechanisms across OpenAI’s data center and compute infrastructure portfolio. Establish scalable EHS requirements for site de
About the Team At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products. Behind the scenes, making these systems fast, reliable, and cost-efficient requires world-class infrastructure. The Caching Infrastructure team is responsible for building a caching layer that powers many critical use cases at OpenAI. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases. We’re looking for an experienced engineer to help design and scale this critical infrastructure. The ideal candidate has deep experience in distributed caching systems (e.g., Redis, Memcached), networking fundamentals, and Kubernetes-based service orchestration. In This Role, You Will: Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences. Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost. Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs. You Might Thrive In This Role If You: Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems. Have deep expertise with Redis, Memcached, or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning. Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems. Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities. Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence. About OpenAI OpenAI is an AI research and deployment company d
About the Team The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. This role is based in San Francisco, CA. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
Get new networking manager jobs by email
Daily job updates · Unsubscribe anytime