Jobiba hiring network

Rack Power Engineer Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current rack power engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time• Remote
2 days ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through

REMOTEpythonartificial intelligenceai
View job →
O
1mo ago

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

awsrestai
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

pythonci/cdlinux
View job →
N
1mo ago

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We're searching for a highly motivated, technical leader to drive the engineering roadmap and innovation for our rack system software architecture. From firmware, kernel drivers, operating systems, networking, fabrics and associated user mode drivers + manageability software. You will work with component leads internally and engage with industry leading hyperscalar / cloud service providers on taking these products to market. What you’ll be doing: Drive the software end-to-end architecture for NVIDIA's rack-scale products Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization. Ensure high quality & reliable software; serving as a trusted architectural partner to teams requiring

We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne

pythonkubernetesartificial intelligence
View job →

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity As a Graduate Performance Engineer, you will contribute to the design, integration, bring-up, and validation of complex server- and rack-level performance of elaborate, large-scale systems. You will work alongside experienced engineers, software developer, and cross-functional partners while building practical skills in component, system and scale up/out performance optimization and design. Start: September, 2027 Location: Austin, Texas, USA What You’ll Do Support the design, peer review, bring-up, and debug of complex server- and rack-level systems. Assist with modeling, design and evaluation of the complete software and hardware stack identifying performance bottlenecks, possible solutions and testing those outcomes. Collaborate with many different teams across both hardware and software development and testing. What You’ll Bring A bachelor’s or master’s degree in electrical engineering, computer engineering, or a related discipline, completed before the role’s start date. Equivalent relevant education or practical experience will also be considered. Foundational knowledge of software development, hardware architecture and general understanding of performance implications. Hands-on experience gained through coursework, laboratories, internships, research, student projects, or personal projects. Ability to analyze technical problems, document your work, communicate clearly, and collaborate effectively. Curiosity, sound engineering judgment, a

artificial intelligenceai
View job →

About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world’s most transformative technologies. Our new AI Engineering Campus in Austin will play a central role in building the future of AI computing. The Opportunity We are seeking a system validation engineering intern to help drive server blade and rack validation efforts for next-generation AI infrastructure hardware systems. This role focuses on post-silicon system validation across the full lifecycle of server hardware systems, ensuring functional and performance meets product objectives. You will help drive end-to-end blade and rack validation including development, execution, and debug while collaborating across silicon, firmware, systems, and platform teams. The Blade and Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Type: 12-week summer internship Timing: May - August (exact dates to be confirmed) Commitment: Full-time What You’ll Do Help drive and execute post-silicon validation goals of AI compute blades and racks including testcase planning, development, and automation Help drive validation testcase execution and system debug against program achievements and report validation progress and risks. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness Triage test failures, collect debug data, and collaborate on root cause analysis. Track validation coverage and continuously improve test processes and infrastructure. What You’ll Bring Working towards a Bachelor's

pythonlinuxartificial intelligence
View job →
G
4 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonartificial intelligenceai
View job →
G
4 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonartificial intelligenceai
View job →
G
8 days ago

About us Graphcore is one of the world’s leading innovators in artificial intelligence compute. We are developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and support the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a family of companies responsible for some of the world’s most transformative technologies. Together, we share a bold vision to enable advanced artificial intelligence and ensure its benefits are accessible to everyone. Graphcore brings together AI researchers, silicon designers, software engineers and systems architects to solve complex technical challenges and deliver innovative computing solutions. Job Summary The Principal Electrical Engineer will be a technical authority within Data Center Engineering, leading the architecture and delivery of safe, resilient and scalable electrical infrastructure for high-density AI computing environments. Working with internal teams, data center developers, utilities, consultants and equipment partners, this role will guide projects from early technical studies through design, construction, commissioning, operation and lifecycle improvement. The successful candidate must reside in, or be willing to relocate to, Austin, Texas. Approximately 10% travel may be required. The Team The Data Center Engineering team is responsible for defining and enabling the infrastructure needed to deploy and operate Graphcore’s computing systems at scale. The team works across electrical, mechanical, thermal, controls, systems and operational disciplines, collaborating with external engineering and construction partners to deliver reliable, efficient and maintainable data center environments. Responsibilities and Duties Act as the technical authority for electrical engineering across data center infrastructure projects, from the utility or on-site power source through to the IT rack. Lead electrical archit

aiexcelsem
View job →
G
Graphcore
📍 Austin• Full-time
8 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonaiexcel
View job →
G
Graphcore
📍 Austin• Full-time
8 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonaiexcel
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for signal integrity (SI) system design engineers who have a deep expertise in the SI area, and hold strong system level design knowledge This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead system signal integrity (SI) design for AI supercomputer product in the data center application. Collaborate with chip, package, boards, rack and system engineers, design partners to drive system SI design and develop innovative interconnect and high-speed technologies Identify and evaluate new technologies and methodologies to improve signal and power integrity in product design, and contribute to the development of new products and technology by providing expertise in signal integrity Perform simulation and modeling to identify and troubleshoot signal integrity issues Lead system interconnect design, bring up and qualification As the scope of the role and team grows, understand and influence roadmaps for hardware partners for our datacenter networks, racks, and buildings. You might thrive in this role if you: Have at least 10 years of industry experience, including experience design hardware system and SerDes testing for data center applications Have a strong bias toward action, and won’t take no for an answer. Have experience and good knowledge of system design experience in the SI areas, from chip, SerDes, board, rack level Have ex

awsrestai
View job →
G
1 day ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonartificial intelligenceai
View job →

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

pythonaiexcel
View job →
🔔

Get new rack power engineer jobs by email

Daily job updates · Unsubscribe anytime