Applying for Principal Firmware Engineer - Data Center Server Management at Nvidia? Compare this live job with your resume first.

🎯 Tailor my resume
←Jobiba.Me
N
Active1w ago

Principal Firmware Engineer - Data Center Server Management

Nvidia·📍 Ca, Santa Clara, United States

Employment

Full-Time

Work mode

On-site

Experience

SENIOR LEVEL

Salary

Not disclosed

Salary not disclosed

Check market pay for comparable Principal Firmware Engineer roles before applying.

Salary →

Role overview

Job description

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. NVIDIA GH200 superchip provides performance and productivity required for strong scaling for HPC and generative AI workload. Scale out is inherent to design of this massive superchip. We are looking for expert engineers to come and help design rack level solutions for next generation scaling AI supercomputing platforms.

We are looking for a strong technical architect to own end to end manageability architecture for these products in data centers. You will work with various component leads internally and externally, drive customer use cases, align architecture with customer requirements and release best products to market.

Join us at the forefront of technological advancement.

What you’ll be doing:

  • Drive server management for large clusters and data centers deploying GPUs and Grace solution from Nvidia.

  • Work with data center architects and cloud customers to narrow down on requirements for implementation to ensure speed of light product development.

  • Work with internal teams to make sure requirements are designed and implemented in right way with each firmware and software module

  • Collaborate with other leads to design & build data center health management workflow.

  • Drive reliability and optimization in firmware architecture from a data center view point.

  • Work closely with cluster bring up team and resolve issues at Speed of Light

  • Own firmware delivered to data centers in terms of quality, reliability and telemetry performance.

What we need to see:

  • 15+ years of relevant experience working on server firmware (BMC) and platform software development

  • BS, MS, or PhD in EE/CS or related field of education or equivalent experience

  • Hands on experience with data center health management workflow. Proven record of delivering server firmware for large data centers..

  • Strong knowledge of data center management, server architecture and server manageability in data centers and strong and demonstrable skill in C/C++ and Python

  • Experience programming and debugging skills for server platforms.

  • Experience in SCM (e.g. Git, Perforce) and project management tools like Jira.

  • You should possess excellent written and oral communication skills, good work ethics, high sense of team-work, love to produce quality work and commitment to finish your tasks every single day.

  • You are a self-starter who loves to find creative solutions to complicated problems and hands on with coding.

Ways to stand out from the crowd:

  • Hands on experience with data center health management

  • Hands on with x86 or ARM system architecture.

  • Proven technical leaders to drive large complex problem with 50+ engineers working

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD for Level 6, and 320,000 USD - 488,750 USD for Level 7.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 28, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

What they are looking for

Skills & requirements

Qualification

What we need to see: 15+ years of relevant experience working on server firmware (BMC) and platform software development BS, MS, or PhD in EE/CS or related field of education or equivalent experience Hands on experience with data center health management workflow

N

Hiring company

Nvidia

Explore this employer's active roles, salary signals and company profile on Jobiba.

Keep exploring

Similar active roles

Fresh roles matched to this title and market.

View all →
R
📍 San Mateo, CA, United States· Full-time

From $345K/yr

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Compute team, you will be the technical anchor for Roblox's GPU and AI accelerator capabilities. This is a battle-tested GPU expert role focused on the machine management layer and above: how GPU hosts are made production-ready, kept healthy, and turned into reliable compute for the workloads that depend on them. You will own the hard problems that show up only at scale, from driver and firmware management to GPU health, reliability, and performance across a rapidly growing fleet of accelerators spanning Roblox data centers and cloud environments. You will set the technical direction for GPU compute and up-level the entire organization's GPU expertise. You will: Serve as the GPU technical leader for the Compute team, partnering across Kubernetes, Machine Bootstrap, Networking, and Cloud to drive GPU strategy end to end. Own the GPU host lifecycle above raw fleet management: driver, firmware, and CUDA stack management, GPU health and telemetry, and remediation of GPU-specific failures (XID errors, ECC, thermal, NVLink and fabric faults). Architect how GPU capacity is exposed to compute platforms, including scheduling, isolation, and integration with Ku

AWSKubernetesGitAI

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets

AWSAzureKubernetesCI/CD

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches

AWSAzureGCPKubernetes
M
📍 San Francisco, California, United States· Full-time

What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set

PythonGitAIC++
H
📍 Texas, United States of America, United States

$147.1K – $230.9K/yr

Principal Embedded Firmware and Software Engineer Description - We are seeking a Principal Embedded Firmware & Software Engineer to lead the design, development, and debugging of embedded software and firmware for computer systems. In this role, you will combine deep, hands-on engineering expertise with system-level technical leadership to ensure seamless integration between software and hardware components, delivering reliable and efficient system performance. You will collaborate closely with cross-functional teams including hardware engineers, software developers, QA, and product managers to bring high-quality products to market. Responsibilities Provide technical leadership for the architecture, development, security, integration, debugging, validation, and deployment of embedded firmware and software including BIOS/UEFI, EFI applications and drivers, embedded controllers, and RTOS-based systems. Analyze hardware and system architectures to define firmware requirements, dependencies, interfaces, integration strategies, and validation approaches. Troubleshoot and resolve firmware issues by designing and implementing enhancements, updates, and programming changes across firmware subsystems. Define and drive firmware integration, verification, and validation strategies, including automated testing, regression testing, and continuous integration . Develop and improve engineering tools and automation using Python and other appropriate technologies for development, debugging, testing, analysis, and validation. Advance CI/CD and DevSecOps practices for embedded development to improve engineering velocity, quality, traceability, and release confidence. Evaluate and apply AI-assisted software development and engineering tools where they can improve developer productivity,

PythonGitLinuxAI
M
📍 Longmont Max Office, United States

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Firmware & Product Test (FPT) team plays a critical role in delivering high-quality enterprise SSD solutions by ensuring firmware functionality, reliability, and compliance. We work across simulation, FPGA, and hardware environments to validate modern storage technologies, build scalable automation, and drive continuous improvement in validation methodologies. Our team values technical excellence, collaboration, and innovation, including the use of AI-enabled tools to enhance engineering productivity and quality. As a Principal Test Development Engineer, you will serve as a technical leader for firmware validation, defining verification strategies, advancing automation frameworks, and driving complex failure analysis efforts. This role offers the opportunity to influence product quality across multiple SSD programs while mentoring engineers and partnering closely with firmware architects to improve testability and validation effectiveness. Responsibilities: Lead verification strategy, test planning, automation, and coverage closure for NVMe front-end firmware features across multiple product lines Architect and enhance scalable Python-based test automation frameworks, CI/CD integration, regression infrastructure, and reporting capabilities Drive root-cause analysis and failure triage using firmware traces, protocol analyzers, system logs, and structured debug methodologies Define validation standards, review test code, mentor engineers, and promote standard methodologies in automation and qua

PythonGitLinuxAI

🔔 Get job alerts

New Principal Firmware Engineer - Data Center Server Management jobs in Ca, Santa Clara, United States, straight to your inbox.

No spam · Unsubscribe anytime