NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
Jobs in United States
Senior System Architect in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current senior system architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA is now looking for a Senior Memory System Engineer to join our ASIC Memory Subsystem team! As a Senior Systems Engineer at NVIDIA, you'll join a group of hardworking engineers to develop and architect innovative Memory Solution for Tegra SoCs. In this position, you'll make a real impact in a multifaceted, technology-focused company. You will work with memory controller/PHY and Platform / System architect, Firmware, SI/PI, Memory suppliers to design and architect cutting edge, high speed and lower power memory technology for NVIDIA CPUs and SOCs. What You Will Be Doing: Analyze future DDR/LPDDR/HBM technologies to determine optimum performance, power, function and RAS in memory for Next generation SOC and Systems. Collaborate with ASIC Architects, Designers, Software and Firmware SW/FW teams to drive memory technology and associated requirements for memory controllers. Define Memory module, Package, and PCB layouts appropriate to the system workloads Debug and bring up memory evaluation / validation and failure issues on memory technology. Collaborate with DRAM suppliers and industry partners on to develop memory and memory related component technology. What We need to see: Bachelor's degree or master’s degree in Electrical Engineering, Computer Engineering (CE), or a related field (or equivalent experience) 10 years of proven track record in DRAM design, module design, or memory sub system design. Deep understanding and strong fundamental of memory design, features, ECC algorithm, SI and PI (Training algorithm) in DDR, LPDDR, and HBM. Strong understanding of memory sub system level interaction with Cache, Memory controller and PHY. Experience in the design, bring-up and validation for memory failure analysis Experience with Python, C/C
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
NVIDIA is seeking a world-class computer architect to contribute to the development of future high-performance computing systems, with a focus on enhancing the power-constrained performance of the hardware. Ideal candidates will have a strong track record of understanding and analyzing memory systems architecture to improve performance per watt (perf/W) and performance per millimeter (perf/mm). A broad perspective across the field of computer architecture and depth in the area of power, performance, and area (PPA) analysis is highly desirable. NVIDIA has pioneered programmable GPUs and the CUDA language and is a world leader in high-performance computing technology, with aggressive plans for future processors. This position offers the opportunity to have a real impact in a fast-moving, technology-focused company. What you will be doing: Develop innovative high-performance processor and system architectures, focusing on the memory system and energy efficiency. Develop architecture and micro-architecture features to improve the state-of-the-art in GPU memory systems, optimizing along the axes of perf/W, perf/mm, and perf/$. Develop and enhance architecture prototype models for power and noise analysis. Participate in performance and power simulation of features to analyze, define, and improve energy per byte. Analyze benchmarks, application workloads, and performance/power simulation and emulation results to identify areas for architecture optimizations. Debug power, performance, and functional issues with high-level models, RTL simulation and emulation, silicon, and systems. Collaborate with outside partners on system infrastructure. What we want to see: 10+ yrs of experience in CPU/GPU architecture, memory systems design with a focus on energy efficiency in the system. Bachelor
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
1671 About the Role We are seeking a Senior Signal Integrity Engineer to develop, validate, and optimize high-speed signaling solutions across blade- and rack-level architectures for advanced compute platforms. This role sits at the intersection of silicon, package, interconnect, board, and system design , with a strong emphasis on hands-on measurement, simulation correlation, and cross-functional technical communication. Key Responsibilities End-to-end signal integrity analysis for blade- and rack-level system architectures. Analyze and optimize high-speed and low-speed I/O interfaces, including PCIe Gen4/5/6, Ethernet, DDR, SerDes, SPI , I2C, etc . and related interconnects. Perform time-domain and frequency-domain simulations using tools such as Ansys HFSS, Keysight ADS, Cadence Sigrity, CST, SPICE , or similar. Support hands-on lab validation using VNA, TDR, BERT, and high-speed oscilloscopes . Correlate simulation results with lab measurements to identify margin gaps, debug issues, and improve design methodology. Collaborate with silicon, package, board, connector, cable, and system design teams to optimize I/O channel performance. Work with interconnect vendors and ODMs to guide board layout, stack-ups, routing rules, and system design decisions. Review schematics, layouts, simulation results, and validation data for blade, backplane, and rack-level hardware. Prepare and communicate clear validation reports, measurement summaries, debug findings, and technical recommendations to internal teams, vendors, and senior technical stakeholders. Required Qualifications Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, or a related field . Strong experience in signal integrity for high-speed digital systems. Hands-on measurement expertise using VNAs, TDRs, BERTs, and high-speed oscilloscopes . Experience measuring and analyzing S-parameters, impedance profiles, eye diagrams, jitter, timing margins, insertion loss, return loss, and cro
Join the NVIDIA's Solutions Engineering team that is reshaping the future of driving! Our goal is to build and deploy scalable solutions for autonomous vehicles and as a result, create safer and more efficient roads. Our team is hands-on, passionate about practical results, and values diversity. You will help craft the application software architecture by working closely with external partners developing on our platform and on collaborations across multiple teams within NVIDIA working on autonomous vehicles. You will also advance and refine the overall drivability of our solution, focusing on integration challenges and using your deep analytical skills to tease through the complexity of the system to find effective solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers, and is committed to fostering a diverse work environment and proud to be an equal opportunity employer. If you are passionate in bringing autonomous vehicles into the world and see the solution come together, we would like to hear from you! What you'll be doing: Shape the application architecture internally, with a focus on perception & sensor fusion, by collaborating closely with architecture and software development teams. Integrate and adapt NVIDIA solutions in target vehicles, ensuring that both perception and sensor fusion are adapted and tuned to meet the desired driving performance and functionality. Lead bring-up activities and provide technical support to resolve functional and perception & sensor fusion related issues. Perform and leverage in-vehicle and simulation test drives for functional and performance analysis on the recorded data. Work with our partners to efficiently integrate hardware and software components, understand the system architecture, profile performance, identify bottlenecks, and drive optimization <
Senior Systems Engineer (Flight Crew Operations Integrator) Company: The Boeing Company Boeing Defense, Space & Security (BDS) is seeking a Senior Systems Engineer (Level 5) to support the Phantom Works organization in Hazelwood, MO . The successful candidates will develop equipment and escape/crew station system integration concepts and other design methods to provide and coordinate product definition for Integrated Product Teams, various engineering functions, production operations, qualification/flight testing, suppliers and external customers throughout the product lifecycle. Position Responsibilities Applies an interdisciplinary, collaborative approach to lead activities to plan, design, develop and verify complex lifecycle balanced system of systems and system solutions. Leads others to evaluate customer/operational needs to define system performance requirements, integrate technical parameters and assure compatibility of all physical, functional and program interfaces. Leads analyses to optimize total system of systems and/or system architecture. Leads analyses for affordability, safety, reliability, maintainability, testability, human systems integration, survivability, vulnerability, susceptibility, system security, regulatory, certification, product assurance and other specialties quality factors into a preferred configuration to ensure mission success. Leads, develops, maintains and identifies improvements the planning, organization, implementation and monitoring of requirements management processes, tools, risk, issues, opportunity management and technology readiness assessment processes. Travel may be required up to 10% of the time; Domestically and/or interna
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, we're not just transforming the world of computer graphics and AI; we're setting the stage for the future of autonomous driving. As a Lead Safety Architect, you will be at the forefront of our autonomous vehicle technology, ensuring its safety at scale. You will collaborate with the most innovative engineers and technologists to integrate safety measures into our latest DRIVE products. This role is paramount in achieving and exceeding NVIDIA's high safety standards, making your work both exciting and impactful! What you’ll be doing: Representing NVIDIA’s functional safety strategy and architectures to the customer Working closely with customers to understand their functional safety requirements and system architectures and feeding those back into the development teams Assisting customers to safely integrate and validate our products in their systems and vehicles Supporting customer facing safety collateral Tailoring functional safety platforms and safety analyses for strategic customers You will be working closely with safety management, solution architects, sales and technical marketing teams to deliver state of the art products
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox is a rich ecosystem encompassing varied product categories across nearly all platforms (mobile, desktop, console, VR, etc.). Keeping our app performant, scalable, and reliable requires a robust technical architecture to empower thousands of engineers to build at scale. As the Senior Product Manager for this space, you’ll drive the creation of foundational systems like our multi-process architecture, resource management systems, and key observability frameworks that power the heart of the Roblox App. You will: Define the capabilities and inner workings of critical systems like our process lifecycle manager, asset download orchestrator and critical observability frameworks like our sessionization system. Drive a multi-year roadmap to modernize our App Architecture across key themes like modularization, network resiliency, and resource efficiency. Dive deep into problem spaces with Engineering to ideate new systems to scale our App. Drive the roadmap for Observability & Logging Systems. Dive deep into monitoring & triage workflows, identifying gaps to create new features and systems that extend our capabilities Empower our customers. Every feature team across Roblox
We are hiring senior engineers to work on the CUDA driver, a core component of our platform for accelerating general purpose computation on the GPU. Our team delivers features and improvements to better realize the potential of NVIDIA hardware for a growing range of computational workloads, ranging from deep learning, scientific computation, and self-driving cars to video games and virtual reality! CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, switches and the operating system. What you'll be doing: As a member of our team, you will use your design abilities, coding expertise, and creativity to deliver the best Compute platform in the world. You will craft elegant solutions to exciting problems and craft the future direction of CUDA as you collaborate with your peers across NVIDIA. You will evangelize, architect, and implement new CUDA features You'll oversee and drive development efforts across multiple teams Collaborate with members of hardware architecture teams Help define forward-looking improvements to the CUDA APIs and programming model Design and maintain performance and precision modeling Write effective, maintainable, and well-tested code Develop code for multiple operating systems What we need to see: Bachelor of Science or Master of Science degree in Computer Science, Electrical Engineering, or related field (or equivalent experience) 15+ years of relevant systems software development experience Strong C programming skills </
Other cities to consider
More places hiring for this role
Get new senior system architect jobs in United States by email
Daily job updates · Unsubscribe anytime