We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
Jobs in United States
Networking Manager in United States
218 active opportunities · Updated October 2026
Showing
15 jobs
Explore current networking manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
$100K – $500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to validate the System Management Controller (SMC) and enable seamless multi-chip integration. In this role, you will design and execute tests, build infrastructure, and debug issues across chiplet-based SoCs. You’ll have the opportunity to work with remote mentorship while contributing to the foundation of scalable multi-die systems. This role is hybrid, based out of Toronto, Ontario, Boston, MA or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proficient in SystemVerilog, SV-UVM, Python, and C/C++ with strong verification skills. Experienced in writing test plans, building infrastructure, and debugging hardware/software flows. Comfortable working with remote mentorship and distributed teams. Familiar with AI-assisted tools like Copilot, Cursor, and Claude to accelerate verification. What We Need Develop and maintain SMC tests and supporting DV infrastructure. Write, execute, and track test plans for chiplet and multi-chip SoC designs. Use C/C++ to develop tests compiled, loaded, and executed directly on the DUT. Triage, analyze, and debug issues in clos
NVIDIA is seeking a strong technology leader to manage our Server Software Technical Program Management (TPM) team. This role is at the cross-section of execution and strategy, leading a team of Senior TPMs who drive the firmware and system software for NVIDIA's next-generation server platforms like DGX, MGX, and HGX. These platforms bring together the full power of NVIDIA GPUs, NVLink, InfiniBand networking, Grace CPUs, and our optimized AI/HPC software stack. This deep technical leadership role focused on the Software Development Processes that brings new server hardware to life. What you'll be doing: Lead a team of TPMs driving the technical software and firmware execution for NVIDIA's NPI (New Product Introduction) and sustaining engineering teams. Drive the end-to-end SDLC for low-level server components, including firmware (BMC, UEFI/BIOS), drivers, and system management software, ensuring alignment with hardware schedules. Collaborate closely with NVIDIA product management and hardware engineering teams to define release plans and program objectives. Build a strong connection and feedback loop between sustaining and NPI engineering teams to improve product quality and development velocity. Lead process improvement initiatives and help propagate SDLC standards across multiple engineering and TPM organizations. You will have the opportunity to interact with diverse technical groups, spanning all organizational levels. What we need to see: Bachelor of Science (or equivalent experience) or Master of Science degree in Computer Science, Electrical Engineering, or related field. 12+ overall years of experience developing and leading complex low-level or system software projects. and 7+ years of experience in a people management role. Deep understanding of system a
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 This role bridges infrastructure and product engineering: you'll build genuine partnerships across product, AI, and enterprise teams so that ownership is shared and velocity is never blocked by platform constraints. As Director, you'll set a forward-looking cloud vision, proactively align with stakeholders across the business, and ensure the platform scales for multi-shard, multi-region growth while meeting security and compliance commitments (SOC 2, business continuity/disaster recovery, and enterprise security frameworks). The Role: Cost Efficiency: Drive significant annual infrastructure savings by migrating data workloads to EKS and self-hosting key services such as OpenSearch. Ingress Convergence: Deprecate legacy frontend ALBs and consolidate to a single EKS-managed ALB per shard, unblocking faster deployments across the org. Coverage & Bench Depth: Eliminate single points of ownership across Networking, OpenSearch, and Terraform through cross-training and targeted hiring into coverage gaps. Stakeholder Alignment: Stand up a recurring alignment cadence with Product, AI, Enterprise, and Security so infra planning is driven by demand, not ad-hoc interrupts. Automation: Ship automated shard buildout via Backstage to remove manual toil from enterprise scaling. Reliability: Cut P0/P1 incidents attributed to Cloud Platform (DNS, ALB misconfiguration) through hardened ingress patterns and Terraform-policy guardrails, including blocking unauthenticated public endpoints. Roadmap Ownership: Deliver a roadmap covering Agent enablement, centralized IaC, and deployment rollout acceleration, tied to AI and
Job Details: Job Description: Join an enthusiastic team of engineers in Intel's Networking Solutions Group (NSG) focused on enabling next generation of programmable Infrastructure Processing Units (IPUs) with our lead customers as part of the Customer Experience Support (CES) organization. Intel brings decades of leadership in networking, virtualization, packet processing, storage, and security to a new class of IPU products that accelerate host networking functions and support emerging use cases such as security, virtualization, storage, load balancing, and data path optimization. Working closely with major cloud service providers and Intel development teams, you will help deliver customized IPU based solutions that enhance isolation, security, performance, storage and system management for our customers. A big part of the day-to-day job is to help customers manage feature request processes, enable solutions, and debug issues. Projects and responsibilities include but are not limited to: • Gain our customers' trust, understand their needs, and build POCs to meet them. Work closely with internal and external partners to understand use cases and requirements. • Be the go-to technical resource for customers building complex Datacenters, AI infrastructure as well as helping them understand performance characteristics for solutions. • Prepare and deliver technical content to customers including presentations, workshops, etc. • Contribute across the full IPU lifecycle, including board and platform bring up, low-level device initialization, OS driver and kernel configuration, system management, feature enablement, use case testing, debugging, and verification. • Defines systems implementation and integration solutions and plans to ensure optimum performance and reliability across hardware, firmware and software w
$42 – $65/hr
Job Title Remote Service Engineer Job Description Serve as the frontline remote technical expert, delivering rapid diagnostics and issue resolution for healthcare customers and field teams while minimizing onsite service through advanced troubleshooting and escalation management. Your role: Provide remote technical support and diagnostics to healthcare customers and field partners using advanced remote access technologies. Respond to customer inquiries as a Philips representative, delivering timely and accurate troubleshooting support. Resolve service issues by leveraging internal tools, systems, and knowledge resources. Escalate complex product issues to higher-level support for further evaluation and resolution. Utilize remote capabilities to efficiently resolve issues and reduce the need for onsite service dispatch. You're the right fit if: You’ve acquired 5+ years of remote software system experience along with a solid background in structured query language (SQL) and Windows server environments. Your skills include computer networking, technical solutions authoring, and the ability to solve mechanical and electrical issues. You have an associates degree or higher, technical discipline preferred. You must be able to successfully perform the following minimum Physical, Cognitive and Environmental job requirements with or without accommodation for this Field Service position. You’re a team player who is versatile and keen to work with a variety of Philips products. How we work together We believe that we are better together than apart. For
The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. Roles and Responsibilities: Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti
About the Team OpenAI's Industrial Compute organization is building and scaling the infrastructure required to support frontier AI. The Infrastructure Strategic Sourcing team connects technical and project requirements to supplier readiness, contracting, purchasing, equipment delivery, and portfolio-level risk visibility across owner-furnished contractor-installed equipment (OFCI), data center networking, rack systems and integration, fiber, cabling, optical interconnects, and related infrastructure. The team partners across Pre-Construction, Design, Construction, Electrical and Mechanical Engineering, Network Engineering, Hardware and Rack Delivery, Strategic Sourcing, Procurement, Legal, Finance, Accounts Payable, Logistics, and external suppliers. We build the operating mechanisms that keep sourcing decisions, purchase execution, long-lead equipment, network and fiber dependencies, rack readiness, and delivery commitments aligned to infrastructure schedules. About the Role We are seeking an Infrastructure Sourcing Operations Lead to own procurement operations across pre-construction, design, construction, and sourcing through purchase order issuance, while maintaining visibility through invoice resolution, production, logistics, delivery, installation, and readiness. The portfolio includes electrical and mechanical OFCI, networking equipment, rack systems and integration, fiber, cabling, optical interconnects, and other infrastructure required to bring capacity online. In this role, you will set priorities, make or escalate decisions that affect cost, supplier relationships, contractual position, and delivery schedules, and define the standards used by execution support for queue management, documentation, tracker maintenance, and recurring reporting. Success requires sound commercial and program judgment, operational rigor, systems thinking, and the ability to turn incomplete information across vendors, tools, and project teams into clear decisions, accountable
From $127K/yr
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Security Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments—GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Security Software Engineer, you will design and build critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Architect and implement production-grade security services (e.g., auth services, access brokers, secure proxies, key-management infrastructure) that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD. Partner with infrastructure and research engineers to embed security into high-performance compute clusters, enabling rapid model training and deployment without compromising protection. Develop automation and detection tooling to continuously identif
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a Delivery Director, Capacity programs for our on-premises data center builds and neo cloud (GPU cloud) delivery programs. This is a high-visibility, execution-critical role sitting at the intersection of infrastructure engineering, capacity planning, vendor/partner management, and customer delivery. You will own the end-to-end delivery lifecycle for large-scale compute infrastructure — from initial site/capacity commitments through power, networking, and hardware bring-up, to production-ready GPU/compute capacity landing in the hands of internal teams or customers. You'll be the person who turns ambitious infrastructure roadmaps into predictable, on-time, delivery. RESPONSIBILITIES Own delivery of on-prem infrastructure builds — colocation expansions, power/cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing — coordinating across colo providers and partners, network engineering, hardware ops, and vendor teams. Drive neo cloud delivery programs — manage capacity delivery from GPU cloud and neo cloud partners (e.g., colocation/bare-metal/GPU cloud providers), including contract milestones, capacity ramps, SLAs, and go-live readiness. Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power/shell timelines, hardware lead times, logistics, and software/platform readiness into a single critical path.
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role Join our Enterprise Platform team and build the infrastructure foundations that enable the world's largest organizations to run Replit within their security and compliance boundaries. As a Software Engineer on this team, you'll design and implement the deployment flexibility, networking capabilities, authorization systems, and data controls that enterprises require, from single-tenant architectures and private connectivity to custom policy enforcement and customer-managed encryption. You'll work at the intersection of cloud infrastructure and enterprise requirements, partnering with Platform Engineering, Security, and Sales to ship capabilities that unlock adoption at demanding organizations. What You'll Do Build enterprise deployment infrastructure: Design and implement single-tenant and dedicated deployment options, enabling customers to run Replit with the isolation guarantees their security posture requires. Implement private networking capabilities: Build VPC peering, private connectivity, and static IP configurations that allow enterprises to integrate Replit into their existing network architectures. Design authorization services: Build the authorization infrastructure that enforces custom enterprise policies; enabling fine-grained access controls, custom permission models, and policy enforcement that integrates with customers' existing identity and governance systems. Ship data protection features: Implement bring-your-own-key (BYOK) encryption, customer-managed keys, and data residency controls that give enterprises ownership over their most sensitive data. Develop infrastructure automation: Write Terraform modules and automation that enable reliable, repeatable enterprise deployments across reg
About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev
Other cities to consider
More places hiring for this role
Get new networking manager jobs in United States by email
Daily job updates · Unsubscribe anytime