About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
Jobs in United States
Networking Manager in United States
218 active opportunities · Updated October 2026
Showing
15 jobs
Explore current networking manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
About the Team The Core Services team is responsible for building and managing foundational services. It acts as the bridge between core infrastructure (e.g. compute, storage, networking) and product engineering teams, and enables product teams to move fast, build reliably, and scale efficiently. About the Role As a software engineer in the core services team, you will design and operate critical backend platforms such as caching systems, workflow orchestration, metadata stores, and file services. You’ll focus on building highly reliable, scalable, and performant systems that serve as the backbone of our products. We’re looking for people who are passionate about building infrastructure that empowers product teams, love working on distributed systems challenges, and enjoy creating well-designed APIs and abstractions that accelerate development. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain shared infrastructure services such as caching layers, workflow orchestration (Temporal), metadata stores, and file storage services. Collaborate with product teams to provide scalable, reliable primitives that abstract the complexities of distributed systems. Improve performance, resilience, and scalability of core services that power customer-facing applications. You might thrive in this role if you: Have experience with distributed systems, caching infrastructure (e.g., Redis, Memcached), metadata storage (e.g., FoundationDB), or workflow orchestration (e.g., Temporal, Cadence). Have experience running containerized services in cloud environments and integrating them into automated build/test/release (CI/CD) workflows. Understand trade-offs in consistency models, replication strategies, and performance optimization in multi-region systems. Excel at communication and collaboration with cross-functional teams, and are obsesse
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. At Micron Technology, we transform how the world uses information to enrich life for all. The Heterogeneous Integration Group (HIG) HBM Architecture team develops next-generation High-Bandwidth Memory (HBM) solutions that power AI, high-performance computing, cloud infrastructure, and advanced networking systems. The team works across architecture, design, verification, packaging, product engineering, and technology development to evaluate innovative memory architectures and deliver scalable, high-performance semiconductor solutions. As an HBM Design Architect, New College Graduate, you will contribute to the evaluation and development of future HBM and DRAM architectures. Working with experienced architects and engineering teams, you will analyze system and block-level design tradeoffs related to performance, power, area, thermal behavior, reliability, and manufacturability. This role provides an opportunity to leverage AI, Large Language Models (LLMs), and data-driven engineering methodologies to accelerate architecture exploration and improve decision-making. Responsibilities Analyze HBM and DRAM architectures using analytic
$100K – $500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our diverse team of technologists have developed a high performance RISC-V-based CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance teams. This is the plumbing that both teams' performance work stands on — benchmarking automation, data collection, and workload capture across silicon, FPGA/emulation platforms and performance models. You'll work across both teams, enabling performance and software engineers to spend their time on analysis and optimization instead of running experiments by hand. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Great at identifying problems and developing solutions, with a bias toward owning the systems you build. Enjoys building tools and optimizing workflows so other engineers can move faster. Strong Linux systems engineer, c
Linux System Administrator Alexandria, VA Secret clearance or higher The Test and Training Enabling Architecture (TENA) program at Leidos is looking to add a Linux System Administrator . Our mission is to develop tools for the United States Department of Defense test and training communities, to increase the speed of the development cycle, to get capabilities to the warfighter more rapidly. The Linux System Administrator will work within a team engineers, developers and system administrators, adding new capabilities to our enclaves, as well as ensuring the existing systems are secure and performing as intended. This position is required to work on-site in our Alexandria, VA office. Primary Responsibilities: Provide support for installing and maintaining Linux servers. Provide technical leadership for migrating on premise infrastructure to hybrid cloud solution. Provide support for troubleshooting from OS to application-level issues. Manage networking equipment including switches, routers and firewalls. Manage configuration and maintenance of core enclave infrastructure services including DCHP, DNS, e-mail, Apache Web Servers, Tomcat application servers, Atlassian applications, MariaDB database servers. Configure and monitor security tools required for DoD networks. Participate and contribute to design discussions related to improving the capabilities, performance and security posture of new and existing services. Basic Qualifications: US Citizen with at least an active Secret clearance. Bachelor’s degree with 4+ years of experience or a Master’s degree with 2+ years of experience. Additional experience may be considered in lieu of a degree. 4+ years of experience providing System Administration to DoD IT systems. DoD 8570 IAT level I
$100K – $500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is looking for a mid- to senior-level Physical Design Engineer who will contribute to the physical design of high-performance chips for industry-leading AI/ML architectures, spanning implementation from synthesis through tapeout. You will partner with front-end and physical design engineers to optimize floorplanning, timing, power, performance, and area across multiple IPs. Along the way, you will build end-to-end ASIC expertise while learning from experienced engineers across the chip development process. This role is hybrid , based out of Austin, TX, Fort Collins, CO, or Santa Clara, CA . We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are An engineer excited to work on high-performance designs for industry-leading AI/ML architectures. A collaborative problem solver who enjoys working with experienced engineers across ASIC disciplines. Grounded in logic design fundamentals and gate- and transistor-level implementation. Curious about how early architectural and RTL decisions shape physical implementation and final chip quality. What We Need A BS, MS, or PhD in Electrical Engineering, Computer Engineering, Computer Science, or a relat
$100K – $500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is looking for a mid- to senior-level Physical Design Engineer with a strong background in static timing analysis (STA). In this role, you will help converge multi-million-gate designs on advanced process nodes and partner with cross-functional teams through final signoff and tapeout. This role is hybrid , based out of Austin, TX, Fort Collins, CO, or Santa Clara, CA . We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A physical design engineer with 8+ years of experience in STA, timing closure, and advanced-node design implementation. Experienced with timing flows, constraints, clocking, CDC, RDC, derates, uncertainties, guardbanding, and final signoff. Skilled in writing and debugging SDC constraints, creating ECOs, developing closure strategies, and analyzing worst-case corners. A strong collaborator and communicator who can drive results across distributed, cross-functional engineering teams. What We Need Drive timing convergence across blocks, partitions, sections, and top-level designs; generate and validate timing models and publish results. Analyze timing violations, develop and propagate fixes, debug timing-flow issues, and
Your challenge As a Sales Engineer, you apply your project experience and industry knowledge to devise engaging industry-specific demos and proofs of concept. You support the Business Development team as an industry advisor during the sales process and work with the Delivery Advisory team on advisory assignments for your industry. You are responsible for: Understanding and researching the competitive differentiators and unique selling points of our targeted industry solutions. Developing, fine-tuning, and managing our unique selling points and storylines for a range of products in one or more of our core industries . Showcasing your solution and industry knowledge at commercial conferences through networking and product demonstrations. Developing a set of industry-specific product demos based on your understanding of their key differentiators. Sharing input and feedback from prospects, customers, and the market at large with the Standard Demo team. Expanding the functionalities of the standard demo in coordination with the Standard Demo team. Engaging and convincing prospects and customers by presenting demos and starting the conversation. Translating customer-specific requirements into high level solution designs in close collaboration with customers and our internal competence centers. Documenting the proposed solution’s scope and project approach. Supporting the advisory team in advisory assignments within your industry. Assisting in the formulation of RFx responses with solution and industry expertise and providing marketing support. Acting as a trusted advisor towards our prospects and customers. About you Essential talents and qualifications: A Master’s or Bachelor’s degree in Operations Research, Mathematics, Programming, Supply Chain, or similar. 3-5 years of experience in supporting sales cycles. In-depth knowledge of supply chain planning processes and solutions. Advanced industry knowledge. Experience in B2B sales, ideally in enterprise soft
Your challenge As a Sales Engineer, you apply your project experience and industry knowledge to devise engaging industry-specific demos and proofs of concept. You support the Business Development team as an industry advisor during the sales process and work with the Delivery Advisory team on advisory assignments for your industry. You are responsible for: Understanding and researching the competitive differentiators and unique selling points of our targeted industry solutions. Developing, fine-tuning, and managing our unique selling points and storylines for a range of products in one or more of our core industries . Showcasing your solution and industry knowledge at commercial conferences through networking and product demonstrations. Developing a set of industry-specific product demos based on your understanding of their key differentiators. Sharing input and feedback from prospects, customers, and the market at large with the Standard Demo team. Expanding the functionalities of the standard demo in coordination with the Standard Demo team. Engaging and convincing prospects and customers by presenting demos and starting the conversation. Translating customer-specific requirements into high level solution designs in close collaboration with customers and our internal competence centers. Documenting the proposed solution’s scope and project approach. Supporting the advisory team in advisory assignments within your industry. Assisting in the formulation of RFx responses with solution and industry expertise and providing marketing support. Acting as a trusted advisor towards our prospects and customers. About you Essential talents and qualifications: A Master’s or Bachelor’s degree in Operations Research, Mathematics, Programming, Supply Chain, or similar. 3-5 years of experience in supporting sales cycles. In-depth knowledge of supply chain planning processes and solutions. Advanced industry knowledge. Experience in B2B sales, ideally in enterprise soft
C$100K – C$500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking an Physical Design Engineer to lead cross-functional efforts to solve complex physical design challenges and develop end-to-end RTL-to-GDS methodologies across advanced nodes, with a strong focus on PPA and runtime improvements. The engineer will architect, integrate, and deploy AI/ML-driven solutions into production physical design flows, creating custom CAD tools and partnering with internal teams and EDA vendors to drive next-generation, ML-enabled capabilities. This role is hybrid, based out of Santa Clara, CA or Austin, TX or Fort Collins, CO. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who you are BS in Electrical or Computer Engineering (or equivalent experience) with 5+ years in Physical Design CAD methodology at advanced nodes. Proven track record improving PPA and/or runtime on high-performance, low-power taped-out designs. Hands-on with industry-standard EDA tools (e.g., Fusion Compiler) across synthesis, P&R, STA, signoff, and hierarchical flows. Strong Python/Tcl and data skills, with interest or experience in ML frameworks (PyTorch, TensorFlow), and the ability to drive complex projects independent
$100K – $500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to join our CPU design team to define and implement RTL for high-performance CPUs. You’ll work on a CPU based on RISC-V ISA, collaborating with DV, PD, and performance teams to deliver a functional, timing, and power-converged design. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in CPU microarchitecture with expertise in Rename, Scheduler, ROB, Load Store, Branch Prediction, Cache or Datapath. Skilled in RTL coding (Verilog/VHDL) and familiar with industry-standard tools for simulation, synthesis, and power analysis. Proficient in debugging RTL/logic across multiple design hierarchies and pre/post-silicon environments. Background in microarchitecture definition, design specification, and performance-driven trade-off analysis. What We Need Own RTL design and microarchitecture development for a portion of a CPU block of a high-performance RISC-V CPU. Collaborate closely with DV, PD, and performance engineers to meet functional, timing, and power goals. Use innovative techniques to optimize power, performance, and
$100K – $500K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking talented Physical Design Engineers to implement high-performance blocks for our industry-leading CPU and AI/ML architectures. You'll own the complete implementation flow from synthesis to tapeout, working alongside world-class engineers to push the boundaries of performance, power, and area. If you're passionate about crafting silicon that powers the future of AI computing and thrive on solving complex design challenges, we want you on our team. This role is hybrid , based out of Austin, TX, Santa Clara, CA or Fort Collins, CO. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A hands-on engineer with deep expertise in SOC/ASIC physical design and a track record of successful tapeouts. Passionate about optimizing PPA through innovative implementation techniques and close RTL collaboration. Strong problem solver who excels at debugging complex issues across design hierarchies. Collaborative team player who thrives in fast-paced, technically challenging environments. What We Need BS/MS/PhD in EE/ECE/CE/CS with proven experience in synthesis, PnR, and timing closure on taped-out designs. Expertise with industry-
Other cities to consider
More places hiring for this role
Get new networking manager jobs in United States by email
Daily job updates · Unsubscribe anytime