Jobs in United States

Principal System Power Management And Performance Architect in United States

334 active opportunities · Updated October 2026

Explore current principal system power management and performance architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $154K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

PythonKubernetesAISwift
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads & Discovery business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver more value to our users and our advertisers. As a Machine Learning Infrastructure Engineer, you’ll build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You will: You will co-design models and systems, working at the intersection of model architecture and ML infrastructure, partnering closely with core modelers, data and AI infrastructure engineers, and product teams to push the boundaries of large-scale training and serving. Your work will span recommendation, search, and agentic applications, including large transformer architectures, LLMs, generative rankers, and efficient offline and online content-understanding systems. You will investigate model, data, and systems tradeoffs end to end—from data pipelines and distributed training to low-latency inference and production serving. This includes designing efficient KV-cache strategies, applying p

AWSGitMachine LearningAI
P
📍 United States· Full-time
✓ High-confidence listingCompany trend -85.6%

From $285.5K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business-critical online systems at massive scale, driving the reliability, efficiency, and evolution behind every core Pinner and Advertiser experience. You'll lead major efforts like multi-region deployment and Kubernetes migration, set the standard for operational excellence, and define the long-term vision for our online serving infrastructure, supporting machine learning and product innovation across the company. This is an opportunity for high-impact technical leadership, broad visibility, and cross-functional influence at the heart of Pinterest’s platform. What you’ll do: Improve reliability, scalability and infra efficiency for Pinterest’s critical online systems across storage and caching, online service and realtime analytics syste

PythonJavaAWSKubernetes
M
📍 O Fallon, Missouri, United States
✓ High-confidence listingCompany trend +212.5%
Quick readStrong listing-quality and freshness signals

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Principal Software Engineer Principal Software Engineer Position Overview The Principal Software Engineer is a senior individual contributor within Architecture & Technology (A&T), responsible for driving enterprise engineering strategy, architecture standards, and technology excellence across Mastercard. This role provides deep technical leadership across distributed systems, cloud-native platforms, resiliency, observability, and software engineering practices while influencing technology direction across multiple teams and domains. The Principal Engineer partners with senior engineering leaders, architects, and platform organizations to define architectural standards, establish reusable patterns, and guide critical technology decisions. Through technical expertise, thought leadership, and cross-functional influence, this role helps teams build secure, scalable, reliable, and operationally excellent solutions that align with Mastercard's long-term engineering strategy. Role • Provide technical leadership and architectural guidance across multiple engineering organizations, driving consistent adoption of engineering standards, best practices, and enterprise technology patterns. • Partner with platform CTOs, architects, and engineering leaders to evaluate technology investments, transformation initi

AWSAzureGCPAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that our model's training environment is as realistic as possible. This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness. About the Role We're looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You'll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments. This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality. In this role, you will Design, build, and evolve the integration between the Codex harness that powers OpenAI's products and research training infrastructure used for training GPT-next Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance Bu

PythonAWSRestAI
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build Core Data Engineering Primitives at Cloud Scale Data pipelines are foundational infrastructure — when they're fast, correct, and maintainable, customers build on them with confidence. If you've spent the bulk of your career building large-scale data infrastructure — designing streaming or transformation primitives, reasoning hard about consistency and fault tolerance, and owning the systems that run under millions of customer workloads — this role might be for you. You'll be working on the streaming and transformation layer at Snowflake: the constructs that define how customers move, shape, and maintain data. AI has a real presence in this work — in how customers use these pipelines and in how we think about building them — but the core job is hard distributed systems engineering, and that's what we're hiring for. About the Team We build the core data engineering primitives that power Snowflake's streaming and transformation capabilities. From the constructs customers use to define real-time pipelines to the execution fabric that makes those pipelines reliable and cost-efficient at cloud scale, our team owns the full stack of declarative data engineering. We're a small, high-ownership team operating close to the product — which means your decisions ship, your architec

JavaAIC++Go
C
📍 Austin, TX, United States· Full-time
✓ Quality checkedCompany trend -100%

About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or changing a line of code. Internet properties powered by Cloudflare all have web traffic routed through its intelligent global network, which gets smarter with every request. As a result, they see significant improvement in performance and a decrease in spam and other attacks. Cloudflare was named to Entrepreneur Magazine’s Top Company Cultures list and ranked among the World’s Most Innovative Companies by Fast Company. At Cloudflare, we’re not looking for people who wait for a polished roadmap; we’re looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you’re the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you’ll fit right in. About the Team Cloudflare handles traffic for almost 25% of the Internet. That’s a lot of data. On the Town Lake team, our mission is to make that data accessible and valuable for users across the company. We connect data from dozens of source systems and make it available so that any user in the company can answer any question in 5 minutes or less, using SQL or plain english. We’re building a modern, agentic-first data lakehouse platform ba

JavaScriptTypeScriptPythonJava
P
📍 San Francisco, CA, United States· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. As a Pinterest employee, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. As a Principal Engineer on the AI Platform team, you'll help architect the infrastructure that powers both Generative AI and Recommender Systems across Pinterest's entire product suite. Our team builds the end-to-end engines for petabyte-scale data orchestration, model training and fine-tuning, and high-performance inference, ensuring our models scale seamlessly to hundreds of millions of inferences per second in service of over 600 million monthly active users.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Safety? At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization’s mission is to become the leader in civil immersive online communities. We systematically detect, remove, and prevent problematic content and behavior, and we make Roblox accounts secure and free from compromise. We cover a broad area of the tech spectrum, including machine learning, classifiers for 3D models, experimentation, automation, detection workflows, and AI-powered text filters. Aligned and partnering with product teams, we use this toolbelt to discover new opportunities, influence and shape the product roadmap and prioritization, build safety products, and measure the impact on our community of users and developers. In doing so, we keep Roblox safe, civil, and inclusive, and we foster positive relationships between people around the world. Why Safety Foundation? Safety Foundation is the infrastructure backbone that powers safety across Roblox. Every piece of content reviewed, every moderation action taken, every safety workflow running anywhere on the platform flows through systems built by this team. We are the platform that other safety teams build on — and right

SQLAWSGitRest
P
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%
Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Millions of people across the world come to Pinterest to find new ideas every day. It’s where they get inspiration, dream about new possibilities and plan for what matters most. Our mission is to help those people find their inspiration and create a life they love. As a Pinterest employee, you’ll be challenged to take on work that upholds this mission and pushes Pinterest forward. As a Principal Engineer on the AI Platform team, you'll help architect the infrastructure that powers both Generative AI and Recommender Systems across Pinterest's entire product suite. Our team builds the end-to-end engines for petabyte-scale data orchestration, model training and fine-tuning, and high-performance inference, ensuring our models scale seamlessly to hundreds of millions of inferences per second in service of over 600 million monthly active users.

JavaAWSRestAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

We are hiring senior engineers to work on the CUDA driver, a core component of our platform for accelerating general purpose computation on the GPU. Our team delivers features and improvements to better realize the potential of NVIDIA hardware for a growing range of computational workloads, ranging from deep learning, scientific computation, and self-driving cars to video games and virtual reality! CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, switches and the operating system. What you'll be doing: As a member of our team, you will use your design abilities, coding expertise, and creativity to deliver the best Compute platform in the world. You will craft elegant solutions to exciting problems and craft the future direction of CUDA as you collaborate with your peers across NVIDIA. You will evangelize, architect, and implement new CUDA features You'll oversee and drive development efforts across multiple teams Collaborate with members of hardware architecture teams Help define forward-looking improvements to the CUDA APIs and programming model Design and maintain performance and precision modeling Write effective, maintainable, and well-tested code Develop code for multiple operating systems What we need to see: Bachelor of Science or Master of Science degree in Computer Science, Electrical Engineering, or related field (or equivalent experience) 15&#43; years of relevant systems software development experience Strong C programming skills </

Artificial IntelligenceAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole — fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements — including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software — ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software — health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior — champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW iss

M
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -100%

What you’ll do Act as the technical lead for large parts of the scanner platform: system architecture, codebase structure, and long-term maintainability. Own core runtime foundations: distributed control, state management, fault handling, and reliability. Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. Build robust observability: logs, metrics, traces, and replayable diagnostics (with privacy constraints). Collaborate with hardware and recon/ML teams to define interfaces, data contracts, timing/synchronization, and failure modes. Lead complex refactors (e.g., message passing / RPC boundaries, modularization, concurrency model) without halting forward progress. What we’re looking for Deep software architecture experience for real-world systems: robotics, instrumentation, medical devices, or other complex distributed products. Strong Python and concurrency background (asyncio, multiprocessing, profiling, performance engineering). Track record of shipping systems that are observable, debuggable, and resilient. Strong technical leadership: clarity, pragmatic trade-offs, and mentoring. Useful experience Building but rock-solid systems: clear interfaces (gRPC/protobuf or equivalent), strong state modeling, and failure handling. High-leverage engineering habits on a lean team: good tests, CI, reproducible dev environments, and fast code review. Practical performance + concurrency work in Python (asyncio, profiling, multiprocessing) and comfort debugging distributed behavior. Security-minded device software: safe defaults, encrypted data paths, and disciplined handling of PII/PHI. Operational thinking: remote updates/management, excellent logging, and diagnostics that make real hardware debuggable.

PythonAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are looking for a 100% hands-on Storage Services Software engineer to join the block storage group. You will be a member of a team that builds the next-generation block storage capabilities and architects a proprietary distributed file system solution from its inception. You will work closely with a variety of teams and architects including the networking team, and external customers. You will take part in defining the software architecture and implementation of the most advanced storage services! Services that will need to meet extreme performance and scalability demands! We have crafted a team of extraordinary people stretching around the globe, whose mission is to push the frontiers of what is possible today and define the platform of tomorrow. At NVIDIA, we work, think and learn as a team. We thrive in a deeply strong environment, and we're passionate about a culture that demands innovation and the highest standards. The rewards are sweet and include collaborating with some of the smartest people in the industry, an aggressive compensation plan that rewards top performers, and the opportunity to work on products that transform the way people work and play. What you’ll be doing: 100% hands-on coding role in C language, Kernel and Userspace Access advanced AI tools and a token budget for code development provided by NVIDIA, the world's AI factory leader. Research, design, implement and test, new and existing, distributed storage services and features of NVIDIA’s block and file storage solution, in both Host and DPU environments. Acquire understanding of the algorithms, the technicalities and the interaction with other components across NVIDIA’s block and file storage ecosystem. Analyze and solve challenging bugs and customer cases in la

L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for a Principal DevOps Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This Principal DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. <l

JavaScriptPythonJavaAWS
🔔

Get new principal system power management and performance architect jobs in United States by email

Daily job updates · Unsubscribe anytime