About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity We are seeking a system validation engineering intern to help drive server blade and rack validation efforts for next-generation AI infrastructure hardware systems. This role focuses on post-silicon system validation across the full lifecycle of server hardware systems, ensuring functional and performance meets product objectives. You will help drive end-to-end blade and rack validation including development, execution, and debug while collaborating across silicon, firmware, systems, and platform teams. The Blade and Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Help drive and execute post-silicon validation goals of AI compute blades and racks including testcase planning, development, and automation Help drive validation testcase execution and system debug against program achievements and report validation progress and risks. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness Triage test failures, collect debug data, and collaborate on root cause analysis. Track validation coverage and continuously improve test processes and infrastructure. What You’ll Bring Working towards a Bachelor's
Jobs in United States
Data Center Controls Network Engineer in United States
2,522 active opportunities · Updated October 2026
Showing
15 jobs
Explore current data center controls network engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity As a Systems Engineering Intern, you will contribute to projects that combine hardware, firmware, and software engineering for advanced AI compute platforms. You will work with experienced engineers on subsystem design, laboratory testing, system validation, automation, and performance analysis. The internship provides hands-on experience with modern hardware development and system-level engineering. You will own clearly defined technical tasks with guidance from the team and document your methods, results, and conclusions. What You Will Do Support the design and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Run laboratory tests and measurements to help evaluate performance, power, signal behavior, and reliability. Contribute to system-level validation by creating scripts and tools that streamline testing, data collection, and analysis. Explore emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and learn how they support advanced computing workloads. Assist with investigations into platform power, cooling, and energy efficiency, including liquid-cooling systems for high-performance processors. Use power meters, oscilloscopes, logic analyzers, or comparable lab equipment under appropriate supervision. Analyze test results, identify unexpected behavior, and work with engineers to reproduce and investigate issues. Collaborate across hardware, firmware, software, m
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the Hardware Platform Development team as a Graduate Systems Engineer. You will contribute to the design, integration, validation, and performance analysis of advanced AI compute platforms. You will work with experienced hardware, firmware, software, mechanical, thermal, and systems engineers throughout the development lifecycle. The role combines subsystem engineering, hands-on laboratory work, test automation, data analysis, troubleshooting, and clear technical documentation. Start: September, 2027 Location: Austin, Texas, USA What You Will Do Contribute to the design, integration, and testing of CPU and high-speed input and output subsystems for advanced compute platforms. Take ownership of defined engineering tasks from requirements and test planning through execution, analysis, and technical review. Develop system-level validation plans, procedures, scripts, and tools that improve test coverage, repeatability, data collection, and analysis. Evaluate platform performance, power, signal behavior, reliability, and interoperability using laboratory measurements and system data. Investigate emerging input and output technologies, including PCIe 6.0 and 800G Ethernet, and assess their use in advanced computing systems. Support platform power, cooling, and energy-efficiency investigations, including liquid-cooling systems for high-performance processor
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity As a Graduate Performance Engineer, you will contribute to the design, integration, bring-up, and validation of complex server- and rack-level performance of elaborate, large-scale systems. You will work alongside experienced engineers, software developer, and cross-functional partners while building practical skills in component, system and scale up/out performance optimization and design. Start: September, 2027 Location: Austin, Texas, USA What You’ll Do Support the design, peer review, bring-up, and debug of complex server- and rack-level systems. Assist with modeling, design and evaluation of the complete software and hardware stack identifying performance bottlenecks, possible solutions and testing those outcomes. Collaborate with many different teams across both hardware and software development and testing. What You’ll Bring A bachelor’s or master’s degree in electrical engineering, computer engineering, or a related discipline, completed before the role’s start date. Equivalent relevant education or practical experience will also be considered. Foundational knowledge of software development, hardware architecture and general understanding of performance implications. Hands-on experience gained through coursework, laboratories, internships, research, student projects, or personal projects. Ability to analyze technical problems, document your work, communicate clearly, and collaborate effectively. Curiosity, sound engineering judgment, a
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that provide the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the hardware platforms that support the next generation of AI systems. The Opportunity We are looking for a recent graduate or early-career engineer to join the Mech/Thermal Platform team as a Graduate Mechanical Engineer. You will contribute to the mechanical design, thermal validation, and verification of advanced AI hardware and data center systems. You will work with experienced mechanical, thermal, hardware, and systems engineers throughout the development lifecycle. The role combines 3D computer-aided design, prototype evaluation, lab testing, data analysis, troubleshooting, and clear engineering documentation. What You Will Do Create and update mechanical parts, assemblies, and drawings using Creo or comparable 3D CAD software. Take ownership of defined mechanical design tasks from requirements and concepts through detailed design, review, release, and verification. Apply mechanical design, materials, manufacturing, heat transfer, and thermodynamics principles to engineering decisions. Develop mechanical and thermal test plans and procedures for prototypes and development systems. Set up and operate laboratory equipment, collect accurate data, analyze results, and document conclusions. Evaluate prototypes and identify mechanical, thermal, assembly, or manufacturability issues. Support design verification, troubleshooting, root-cause analysis, and implementation of verified design improvements. Maintain accurate engineering documentation, including design notes, drawings, test procedures, results, and change r
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity As a Mechanical Engineering Intern, you will contribute to the mechanical design and thermal validation of advanced AI hardware. Working alongside experienced mechanical and thermal engineers, you will gain hands-on experience with 3D computer-aided design (CAD), laboratory setup, test development, prototype evaluation, and engineering documentation. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Contribute to mechanical design activities using Creo or comparable 3D CAD software, including part modeling, assemblies, drawings, and design updates under the guidance of experienced engineers. Work with the thermal engineering team to develop test procedures and help set up laboratory capabilities for evaluating AI hardware. Support thermal and mechanical testing using appropriate laboratory equipment; collect, organize, and analyze test data. Assist the mechanical and thermal design teams with prototype evaluation, troubleshooting, design verification, and documentation of findings. Collaborate with cross-functional engineering partners, communicate progress and issues clearly, and follow applicable laboratory and safety procedures. What You’ll Bring Current enrollment in a bachelor's, master's, or doctoral program in mechanical engineering or a related discipline during the internship. Students in other disciplines with relevant mechanical engineering exper
Job Details: Job Description: Responsible for managing the relationship between Intel Data Center and AI (DCAI) and Dell ISG with focus on driving sales strategy of existing Xeon and AI product offerings while expanding Intel design portfolio within Dell's future roadmap to ensure the acceleration of Xeon and AI adoption. Through the design-in process, demonstrate the skills to showcase a quantifiable preference to Intel technology and integrate key learnings into Intel and Dell marketing organizations. The role is critical to establish great relationships and trust with key individuals at Dell and Intel business units and help both companies understand each other's challenges, opportunities to work towards a solution(s) that allows both companies to succeed. The candidate will have enough technical depth to understand the Intel product roadmap, the customer roadmap needs and what the competition is offering and intelligently recommend how to sell the Intel value proposition. Key Responsibilities include: Serve as sales and technical lead for multiple Dell PowerEdge and AI server programs based on Intel Xeon processors and AI accelerators, and collaborate with marketing team to drive complete end-to-end execution. Partner closely with Intel and Dell engineering, product, marketing, and program teams worldwide to align on features, schedules, quality, and issue resolution, ensuring successful product execution. Drive strategic design-win engagements for future Intel Xeon and AI accelerator products, influencing both Intel and Dell roadmaps through deep customer insight, competitive analysis, and early requirements gathering. Act as the primary interface between Intel and Dell, rapidly resolving business and technical issues, managing escalations, and ensuring the highest levels of cus
Mechanical and Thermal Laboratory Technician Position Summary Graphcore is a globally recognized leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data center hardware that provide the specialized processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Responsibilities The Mechanical and Thermal Laboratory Technician is a hands-on technical role supporting the development and validation of advanced AI hardware systems for data center environments. Working as part of a cross-functional engineering team, this individual will be responsible for executing mechanical and thermal laboratory testing, supporting product validation activities, prototype fabrication and assisting with troubleshooting and root-cause analysis of complex hardware systems. Requirements Associate degree in Mechanical Engineering Technology or a related technical field preferred. Equivalent combinations of education, training, and relevant experience will be considered, including experienced non-degreed candidates or candidates with degrees in unrelated disciplines. Minimum of 5 years of experience working in mechanical laboratories, machine shops, test labs, or similar technical environments. Experience with server hardware platforms and data center equipment. Knowledge of Direct Liquid Cooling (DLC) systems and their implementation in server environments. Experience operating forklifts, pallet jacks, and other material-handling equipment. Key Responsibilities Execute mechanical and thermal test plans to validate hardware designs a
NVIDIA is seeking an a PCB Library Engineer to join our PCB Design Infrastructure team. In this role, you will help develop and maintain the PCB library assets used across NVIDIA's Data Center, AI, Networking, Automotive, and Graphics products. Working alongside experienced PCB designers, library engineers, mechanical engineers, manufacturing engineers, and component engineers, you will create and validate component footprints, schematic symbols, mechanical components, panel definitions, and other critical design assets that enable successful product development. This position provides an excellent opportunity to build expertise in PCB design, manufacturing, component engineering, and design automation while supporting some of the most advanced computing platforms in the world. What you'll be doing: Develop PCB footprints, padstacks, schematic symbols, and mechanical library content using Cadence PCB design tools. Review component datasheets, package drawings, and engineering specifications to create accurate design libraries. Support library verification, release, and documentation processes. Partner with PCB design, mechanical engineering, manufacturing engineering, and operations teams to resolve library-related issues. Learn and apply industry standards, including IPC requirements, DFM, DFA, and DFT principles. Support quality initiatives to ensure library content is accurate, manufacturable, and scalable. Participate in continuous improvement and automation efforts within the library environment. Develop a strong understanding of PCB fabrication, assembly, and component technologies. What We Need to See: BS degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Manufacturing Engineering, or a related field or equivalent experience. <p
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of AI at unprecedented scale. Through a combination of strategic partnerships and self-built campuses, we are developing and operating large-scale data center infrastructure across power, cooling, networking, compute, construction, and site operations. The scale and complexity of this infrastructure introduces a broad range of environmental, health, and safety considerations across site development, design, construction, equipment deployment, commissioning, and ongoing operations. EHS is a critical part of how we build infrastructure that is safe, resilient, compliant, and capable of operating at scale. About the Role We are seeking an EHS Lead to establish and drive environmental, health, and safety strategy across OpenAI’s rapidly expanding compute infrastructure portfolio. This role will develop the EHS framework for large-scale data center development and operations, partnering closely with engineering, construction, infrastructure delivery, facilities, operations, security, legal, environmental, and external development partners. The EHS Lead will help ensure that safety and environmental considerations are embedded into projects from early design and site development through construction, commissioning, and operations. The role will establish standards and operating mechanisms, assess and mitigate risks, oversee EHS performance across internal teams and third-party partners, and provide technical leadership on complex or high-consequence safety issues. Success requires the ability to operate strategically while maintaining strong technical depth and executional rigor in fast-moving, highly complex infrastructure environments. In this role, you will: Develop and own EHS strategy, standards, programs, and operating mechanisms across OpenAI’s data center and compute infrastructure portfolio. Establish scalable EHS requirements for site de
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI's Hardware organization builds supercompute platforms from silicon and boards to full rack-scale systems to power advanced AI workloads. This role owns end-to-end quality for high-speed interconnect hardware across the product lifecycle: early design influence, supplier/contract manufacturer readiness, qualification, ramp, and fleet quality in lab and data center environments. You will be the quality lead for advanced interconnect components and assemblies, including high-speed copper cables, cable cartridges, patch panels, backplane/cable-backplane solutions, high-speed connectors, and related electro-mechanical interfaces. You will partner closely with electrical, mechanical, SI/PI, systems, reliability, operations, and external vendors to prevent escapes and drive rapid, data-driven containment and corrective action. In this role you will: Own quality for advanced interconnect components and assemblies: high-speed connectors, high-speed copper cables, cable cartridges (e.g., cable cassette style assemblies), patch panels & optics, and backplane/cable-backplane interconnect solutions. Drive quality-by-design: participate in design reviews, DFM/DFx, tolerance stacks, material and plating selections, connector mating strategy, strain relief, and assembly methods to reduce variation and field failures. Define and track quality and reliability metrics (DPPM, yield, escapes, RMA/FRACAS trends, Cpk/Ppk where applicable) for interconnects across NPI and m
About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t
Other cities to consider
More places hiring for this role
Get new data center controls network engineer jobs in United States by email
Daily job updates · Unsubscribe anytime