NVIDIA is the world leader in GPU Computing. We are passionate about markets including gaming, automotive, professional vision, HPC, datacenters and networking in addition to our traditional OEM business. NVIDIA is also well positioned as the ‘AI Computing Company’, and NVIDIA GPUs are the brains powering modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We have some of the most experienced and dedicated people in the world working for us. If you are dedicated, forward-thinking, and if working with hard-working technical people across countries sounds exciting, this job is for you. We are now looking for a Software QA Test Development Engineer, you will collaborate with multi-functional groups. SWQA test developer engineer at NVIDIA is responsible for test planning, execution, and reporting, you will also write scripts to automate testing, design and develop tools for QA team, or develop integration tests for validation, so QA engineer can improve productivity or optimize test plan. As a SWQA test developer, you must identify weak spots and constantly design better and creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. What You’ll Be Doing Analyze requirements and design test matrices covering functionality, performance, and edge cases. Develop test plans and cases; build and maintain automated test suites (API / UI / CLI / E2E) in Python. Leverage AI-powered tools and agentic workflows to accelerate test generation, triage, and root cause analysis. Manage the full bug lifecycle — filing, reproduction, and driving multi-functional collaboration to resolution. Reproduce and verify customer-reported issues to ensure quality before release.
Jobiba hiring network
Networking Manager Jobs
555 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current networking manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We are now looking for a Senior Chip Design RTL Design Engineer for the Switch Silicon group. As a Chip Design Engineer at NVIDIA's Networking business unit, you'll join a group of passionate engineers to design and implement the next generation state of the art Switch Silicon chips. In this position, you'll make a real impact in a dynamic, technology-focused company while developing the industry's best high-speed communication devices, delivering the highest throughput and lowest latency! What you'll be doing: Work in a combined design and verification team which develops some of the switch silicon core units. Plan and Design RTL units / blocks according to Arch & Micro arch specifications under challenging constraints with high orientation to power, area, and performance. Build reference models, verify and simulate chip blocks/entities according to specifications. Work closely with multiple teams within organizations such as Architecture, Micro- Architecture, and FW. What we need to see: B.Sc. in Electrical Engineering or Computer Engineering. 4+ years of experience in RTL design or RTL verification. Previous experience in networking - an advantage. A team player with good communication and interpersonal skills. NVIDIA has some of the most forward-thinking and hardworking people in the world working for us. Are you creative and autonomous? Do you love the challenge of crafting the highest performance & lowest power silicon possible? If so, we want to hear from you. Come, join our Switch Silicon design team and help us build the next chip in this exciting and quickly growing field. #LI-Hybrid <
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
About the Team The Core Services team is responsible for building and managing foundational services. It acts as the bridge between core infrastructure (e.g. compute, storage, networking) and product engineering teams, and enables product teams to move fast, build reliably, and scale efficiently. About the Role As a software engineer in the core services team, you will design and operate critical backend platforms such as caching systems, workflow orchestration, metadata stores, and file services. You’ll focus on building highly reliable, scalable, and performant systems that serve as the backbone of our products. We’re looking for people who are passionate about building infrastructure that empowers product teams, love working on distributed systems challenges, and enjoy creating well-designed APIs and abstractions that accelerate development. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain shared infrastructure services such as caching layers, workflow orchestration (Temporal), metadata stores, and file storage services. Collaborate with product teams to provide scalable, reliable primitives that abstract the complexities of distributed systems. Improve performance, resilience, and scalability of core services that power customer-facing applications. You might thrive in this role if you: Have experience with distributed systems, caching infrastructure (e.g., Redis, Memcached), metadata storage (e.g., FoundationDB), or workflow orchestration (e.g., Temporal, Cadence). Have experience running containerized services in cloud environments and integrating them into automated build/test/release (CI/CD) workflows. Understand trade-offs in consistency models, replication strategies, and performance optimization in multi-region systems. Excel at communication and collaboration with cross-functional teams, and are obsesse
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
Build in-demand cloud skills with KP Expert’s Oracle OCI Course Online. Learn Oracle Cloud Infrastructure through hands-on training, real-world projects, and expert mentorship. Master compute, networking, storage, security, and cloud administration while gaining practical experience. Our industry-focused curriculum helps you prepare for Oracle certifications and unlock rewarding cloud career opportunities with confidence. For more information visit: https://www.kpexpert.com/courses/oracle-oci
The roles and responsibilities of a sales executive include: Generating leads through cold calls, emails, and networking Following up with prospects and building strong relationships Presenting and demonstrating products or services to potential customers Negotiating and closing sales deals to meet or exceed targets Maintaining client records and tracking sales progress using CRM tools Providing after-sales support to ensure customer satisfaction
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. At Micron Technology, we transform how the world uses information to enrich life for all. The Heterogeneous Integration Group (HIG) HBM Architecture team develops next-generation High-Bandwidth Memory (HBM) solutions that power AI, high-performance computing, cloud infrastructure, and advanced networking systems. The team works across architecture, design, verification, packaging, product engineering, and technology development to evaluate innovative memory architectures and deliver scalable, high-performance semiconductor solutions. As an HBM Design Architect, New College Graduate, you will contribute to the evaluation and development of future HBM and DRAM architectures. Working with experienced architects and engineering teams, you will analyze system and block-level design tradeoffs related to performance, power, area, thermal behavior, reliability, and manufacturability. This role provides an opportunity to leverage AI, Large Language Models (LLMs), and data-driven engineering methodologies to accelerate architecture exploration and improve decision-making. Responsibilities Analyze HBM and DRAM architectures using analytic
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our diverse team of technologists have developed a high performance RISC-V-based CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance teams. This is the plumbing that both teams' performance work stands on — benchmarking automation, data collection, and workload capture across silicon, FPGA/emulation platforms and performance models. You'll work across both teams, enabling performance and software engineers to spend their time on analysis and optimization instead of running experiments by hand. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Great at identifying problems and developing solutions, with a bias toward owning the systems you build. Enjoys building tools and optimizing workflows so other engineers can move faster. Strong Linux systems engineer, c
Linux System Administrator Alexandria, VA Secret clearance or higher The Test and Training Enabling Architecture (TENA) program at Leidos is looking to add a Linux System Administrator . Our mission is to develop tools for the United States Department of Defense test and training communities, to increase the speed of the development cycle, to get capabilities to the warfighter more rapidly. The Linux System Administrator will work within a team engineers, developers and system administrators, adding new capabilities to our enclaves, as well as ensuring the existing systems are secure and performing as intended. This position is required to work on-site in our Alexandria, VA office. Primary Responsibilities: Provide support for installing and maintaining Linux servers. Provide technical leadership for migrating on premise infrastructure to hybrid cloud solution. Provide support for troubleshooting from OS to application-level issues. Manage networking equipment including switches, routers and firewalls. Manage configuration and maintenance of core enclave infrastructure services including DCHP, DNS, e-mail, Apache Web Servers, Tomcat application servers, Atlassian applications, MariaDB database servers. Configure and monitor security tools required for DoD networks. Participate and contribute to design discussions related to improving the capabilities, performance and security posture of new and existing services. Basic Qualifications: US Citizen with at least an active Secret clearance. Bachelor’s degree with 4+ years of experience or a Master’s degree with 2+ years of experience. Additional experience may be considered in lieu of a degree. 4+ years of experience providing System Administration to DoD IT systems. DoD 8570 IAT level I
Get new networking manager jobs by email
Daily job updates · Unsubscribe anytime