Jobs in United States

Principal System Power Management And Performance Architect in United States

334 active opportunities · Updated October 2026

Explore current principal system power management and performance architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Santa Clara, United States
✓ Quality checkedExact matchCompany trend -8%
Quick readExact title match for your search

NVIDIA builds the silicon behind AI, accelerated computing, and graphics. Every watt of performance and every degree of thermal headroom traces back to decisions made in power, performance, and thermal architecture. We are the Silicon Co-Design Group (SCG). We identify, own, and drive system-level co-design ideas. We start with initial concepts and advance to product differentiation across NVIDIA's roadmap. We are hiring a Principal System Power Management and Performance Architect who operates at the ambiguous boundary where workload behavior, silicon capabilities, firmware policies, and platform constraints collide, and who turns that ambiguity into architecture that survives across multiple silicon generations. SCG scope spans architecture, design, software, operations, platforms, and productization. This role shapes system, platform, and data center features and behavior, and partners with teams across NVIDIA. What You'll Be Doing: The work here is rarely well-defined when it arrives. You will be given problems that appear to be performance gaps or power anomalies and encouraged to build a framework for solving them, not just tackle a single instance. Define the multi-generation roadmap for system-level power and performance features, grounded in prototyping, use-case analysis, and cost/benefit trade-offs across segments. You will decide what problems are worth solving and why. Own the architecture and integration strategy for HSIO power management, DVFS, P-states, and low-power features. Your decisions improve product performance, power, and reliability across product lines — not just the current program. Lead system-level boot and IST architecture defining how power and clock domains initialize, sequence, and recover across complex multi-IP systems where the interaction space is large and the failure modes matter. Drive power management strategy at data

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with engineering teams of key CSP / hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work streams with engineering teams of key CSPs/hyperscale customers to build shared understanding of GPU firmware and system software integration, incorporate their feedback into NVIDIA's feature roadmap and delivery plan, and ensure customer-side automation and recovery procedures are ready before each firmware release. Your cross-CSP visibility enables you to identify patterns in GPU firmware operational challenges that drive systemic improvements no single customer engagement could surface alone. What you'll be doing: Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update sequencing, recovery procedures, and GPU power management Gather and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, security requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into NVIDIA's GPU firmware/software feature roadmap and delivery plan Drive GPU firmware update orchestration for large-scale deployments — multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack Serve as the technical focal point between NVIDIA and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are well-documented and accessible for customer integration Identify cross-CSP GPU SW/FW issue patterns — common update failu

Artificial IntelligenceAI
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are the Snowflake Metadata team. We own Snowflake’s metadata systems that make it easy for customers to query, modify and manage their petabyte-scale data. We develop distributed systems that store and maintain metadata, transaction frameworks that power Snowflake’s query and DML capabilities, asynchronous systems that provide time travel and lifecycle management capabilities and entity metadata supporting DDL capabilities. We also build foundational capabilities that deliver global features like cross-region replication (Snowgrid), data sharing, and data marketplace. AS A PRINCIPAL SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, develop and support fault-tolerant scalable distributed systems for our Snowgrid and Data Sharing teams. Create architecture and design, influence our product roadmap, and take ownership and responsibility over new projects. Analyze fault-tolerance and high availability issues, performance and scale challenges, and solve them. Mentor and grow junior engineers. Understand trade-offs between consistency, performance and costs to build solutions which can meet the demands of rapidly growing services. Ensure operational readiness of

JavaAIGoRust
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $280.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As the Senior/Principal Product Manager for Engine Systems Foundations, you will drive the vision and strategy for the most foundational parts of the Roblox game engine and be hands-on with the execution and delivery of products that impact over 130 million players every day. This team is responsible for the core performance, reliability, and efficiency of the engine, and it owns key features like our memory allocation library, thread/work dispatch system, and the backing APIs that power our creator performance tooling. If you are a visionary product leader who thrives on deeply technical challenges to improve the speed and quality of a system, you’ll be a great fit! The role is based in San Mateo, CA (hybrid with Tues-Thurs onsite). You will: Define the long-term vision and strategy for Systems Foundations, ensuring we have plans in place to continually invest in the core pieces of a high-performance, realtime game engine. Take ownership of the engine-related content in the public Creator Analytics creators use to monitor the experiences on Roblox, ensuring we’re delivering actionable insights. Work with the Creator organization to define and drive end-to-end performance workfl

AWSGitAIC++
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Micron currently has an opening for a Staff or Principal level EUV Equipment Engineer for our Boise, Idaho location. Our Advanced Lithography and Equipment Engineering team enables the technologies that power the future of semiconductor innovation. We work at the leading edge of EUV and High-NA EUV development, partnering across process, manufacturing, facilities, and supplier organizations to deliver world-class equipment performance and support Micron’s technology roadmap. This EUV Equipment Engineer is a technical leadership role focused on maximizing the performance, reliability, and capability of EUV lithography systems. This position plays a critical role in advancing next-generation semiconductor technologies through equipment optimization, problem solving, supplier engagement, and collaboration across engineering disciplines. Responsibilities: Own the performance, reliability, availability, and productivity of assigned EUV lithography equipment Lead equipment installations, upgrades, qualifications, maintenance strategies, and continuous improvement initiatives Drive improvements in scanner performance, tool matching, imaging stability, defectivity reduction, contamination control, and recipe optimization Analyze equipment data and establish monitoring systems to improve overlay, focus, source performance, automation, and overall equipment health Provide technical leadership through supplier management, cross-functional collaboration, mentoring, a

AIRecruitment
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives, spanning AI research specialists, silicon designers, software engineers and systems architects. Job Summary We are looking for an experienced Principal Engineer to join our System Management team and help lead the development of critical interfaces used by internal and external customers to manage system state. You will provide technical leadership within assigned areas of System Management, guide architecture and implementation choices, mentor engineers and translate broader technical direction into effective execution. This is a hands-on engineering role for someone who can lead complex technical work, improve reliability and operational readiness, and collaborate effectively across multiple engineering disciplines. The Team The System Management team sits within the Software Platform group and helps build Graphcore products into large-scale AI solutions for our customers. The team is responsible for developing the interfaces between hardware, AI software and frameworks, as well as providing interfaces for public and private cloud environments. This includes system management capabilities that abstract complex hardware administration and enable reliable deployment and operation at scale. As one of the first teams to work with new hardware and software, we regularly solve complex system-level problems

PythonKubernetesCI/CDGit
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. NVIDIA’s Silicon Co-Design Group is hiring a Chip Lead to serve as the technical lead for one of our most consequential silicon programs! This is not project management, and it is not a senior IC role. You are the person the program partners with on its hardest technical questions — when the right answer is not obvious, and when leadership needs a single technical perspective to align on direction. You are accountable for the technical integrity of the chip end-to-end. You guide co-design feature integration, help resolve the toughest multi-functional bugs, and serve as the project-specific custodian of the qualification playbook. The program runs cleaner because you are on it! What you’ll be doing: You will partner across design, validation, software, and manufacturing to keep the program’s technical narrative clear and on track. Day to day, you will: Serve as the single technical point of contact for multi-functional decisions, issues, and trade-offs. Co-lead program-level feature integration from chip to system, surfacing inter-function dependencies and guiding them to resolution. Help resolve the program’s hardest multi-functional bugs by translating ambiguous, multi-team symptoms into root-cause closure on areas such as HBM, power and thermal, high-speed I/O, and packaging. Steward the qualification playbook. When the playbook does not fit a situation, guide the mitigations and capture the lessons as reusable methodology for other SSG programs. Shape the program’s technical narrative by surfacing key risks, trade

AIProject Management
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $280.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the Role Inference Platform is Roblox's multi-tier orchestration system for microservices, powering AI Inference across Roblox. Through a simple developer API, we deliver a "deploy and forget" runtime, hiding the complexity of scheduling, scaling, and reliably running services across our on-prem and multi-cloud footprint at global scale. As Principal Product Manager, Jobs Platform, you'll take the helm at a defining moment - leading the charge as we scale the platform to become the default runtime for critical Roblox services worldwide. You Will Own Inference Platform end-to-end - set the multi-year vision for how Roblox engineers deploy, run, and scale AI and other services across our Core and Edge Datacenters, and cloud. Power Roblox's AI future - build the platform that brings frontier models and next-gen AI workloads to life, with the primitives, scheduling guarantees, and resource classes AI teams need to move fast. Evolve the platform's techni

ReactAWSAzureGCP
O
📍 United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches

AWSAzureGCPKubernetes
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA designs the silicon behind AI, accelerated computing, and graphics. Power and thermal architecture decisions sit behind every watt of performance and every degree of thermal headroom! We are the Silicon Co-Design Group (SCG). We are hiring a Principal Architect to scout the research and industry landscape, identify emerging system-level co-design ideas in power and thermal, and drive them across teams into product differentiation across NVIDIA's roadmap. This role shapes system, platform, and datacenter feature/behavior, and partners with teams across Nvidia. SCG scope spans across architecture, design, software, operations, platforms, and productization. What you'll be doing: Architect next-generation system, platform, and datacenter-level power and thermal co-design solutions. Scan internal research, academia, standards bodies, and silicon, memory, packaging, and platform partners for what is emerging. Build the product differentiation case for each candidate idea, performance, power, reliability, schedule, cost, and brainstorm what is worth pursuing. Lead end-to-end co-design from concept to product. Drive alignment across architecture, VLSI, softw

Machine LearningAI
N
📍 Remote, United States· Remote
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As a Principal Engineer for Security Architecture, within our Security Engineering organization, you will own a core security domain of the AI factory: the architecture, the paved road that delivers it, and much of the code underneath. You will hold the security design bar across DGX Cloud from inside the teams doing the building, and this is a founding seat on a new team. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and Security Architecture is the function inside it that holds the design bar. Security here is fleet horizontal and stack vertical, so your work will cross every DGX Cloud engineering organization: you will embed with the teams building GPU clusters, control planes, and services, join their designs as a participant rather than an approver, and leave behind systems in which an entire class of risk is no longer possible. There is no architecture review board here and no approval queue. You are a senior IC with deep security domain knowledge, and the security bar holds because you helped set it and then helped ship it. What You Will Be Doing: Own a Security Domain End to End: Take architectural ownership of a core domain of DGX Cloud security, from the design through the system running in production. That could be tenant and GPU workload isolation, workload identity, infrastructure and network, supply-chain provenance, hardened baselines and patching, or deploy-time policy and admission control. Embed with the Teams Building It: Join the design early, write the code, and help land it. The posture is not "you did this wrong." It is "here are the considerations we need to meet, I will help, let's go to work." Build Paved Roads, Not

KubernetesLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

AIGoExcelSEM
MT
📍 Richardson, TX, United States
✓ Quality checkedCompany trend +1266.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. We are seeking a Principal Analog Design Engineer to lead the design, integration, and delivery of advanced analog and mixed-signal IPs for High Bandwidth Memory (HBM) products. This role is critical in developing and integrating high-performance analog subsystems within the HBM logic die, enabling industry-leading bandwidth, power efficiency, and reliability. As a principal engineer, you will provide deep technical leadership across analog design, IP integration, system alignment, and silicon execution, driving end-to-end success of HBM solutions. Responsibilities will include, but are not limited to: Lead the d esign and own critical HBM analog circuits, including: High ‑ speed transmitters and receivers Clock generation and distribution (PLLs, DLLs, CDRs) SerDes ‑ related analog blocks Biasing, reference, and calibration circuits

AIRecruitment
F
📍 San Francisco, California, United States
✓ High-confidence listingCompany trend +100%
Quick readStrong listing-quality and freshness signals

Fin , now part of Salesforce, is on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey, from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk, giving modern support teams one single system. Together with Salesforce, the #1 AI CRM, where humans with agents drive customer success, we're building the future of customer experience. Here, ambition meets action. Tech meets trust. And innovation isn't a buzzword, it's a way of life. The world of work as we know it is changing, and we're looking for Trailblazers who are passionate about bettering business and the world through AI. Ready to level up your career at the company leading workforce transformation in the agentic era? You're in the right place. Agentforce is the future of AI, and you are the future of Salesforce. What's the opportunity? Solutions PMM at Fin owns the narratives that shape how Fin wins in market: the stories the field uses in customer conversations, the narratives that power our campaigns, the sales plays that power pipeline, and the competitive strategy that helps us take and defend ground. This is a senior IC role. You drive impact through the quality of your thinking and your ability to move sales, enablement, and campaigns to execute with confidence. This role focuses that mission on who we sell to and how we reach them: our priority industries, our key segments - starting with Enterprise - and the campaign messaging that carries it all to market. It's where audience insight becomes a bill of materials the field and demand gen can run with. As Fin scales beyond a one-size-fits-all story, this is how we get specifi

AISalesforceCustomer ServiceCRM
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore’s AI infrastructure platforms. This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters. The Team Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. The Systems Engineering and Platform Validation team ensures Graphcore’s AI compute platforms are reliable, diagnosable, and operationally robust at scale. The team co

PythonLinuxAIC++
🔔

Get new principal system power management and performance architect jobs in United States by email

Daily job updates · Unsubscribe anytime