What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. As a Fleet Engineer, you'll own the reliability and continuous improvement of our deployed robotic fleet — leading hands-on investigations into how and why robots fail in the field, across the mobile base, charging/docking, motion and power, connectivity (modem), and sensor hardware. You'll combine remote data analysis with bench/lab failure analysis at our Austin HQ, turning field-technician reports and fleet data into clear problem statements, validated root causes, and corrective actions driven to closure with engineering, operations, manufacturing, and vendors. We are hiring a Lead Engineer, Issue Management & Triage to lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering. This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet. You will work deeply with Engineering teams to design classification frameworks, build internal tools, and develop automation pipelines that improve reliability at scale. Location: Austin preferred, Remote possible (U.S.) Travel: if remote up to ~50% travel to Austin, TX (especially in the your first 90 days) What You’ll Do: Own Issue Management & Triage Systems Design and own end-to-end systems for issue intake, triage, and escalation. Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization. Build Tools & Automation (Hands-On) Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort. Contribute directly to codebases (Python, backend services) and partner with Engineer
Jobs in United States
Senior System And Manufacturing Codesign Architect in United States
1,941 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior system and manufacturing codesign architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. About our Team: Micron’s Industrial and Physical AI team is driving the transformation of semiconductor manufacturing through Autonomous Operations, AI, robotics, and digital twin technologies! We develop and deploy innovative solutions across Micron’s global fabrication and assembly/test facilities, enabling smarter, safer, and more efficient operations at scale. Position Overview: We are seeking a hands-on Full-Stack AI Engineer to design, build, and deploy production-grade AI applications that support Micron's Autonomous Operations initiatives. This role owns the end-to-end development lifecycle, from data pipelines and AI models to APIs, web applications, digital twin integrations, and cloud/edge deployments, delivering impactful solutions for engineers, operators, and business leaders worldwide. Responsibilities: Design, architect, and deliver end-to-end AI products, including data ingestion pipelines, feature engineering, model training/inference, APIs, user interfaces, and application monitoring. Build and maintain modern front-end applications using React, Angular, or Streamlit, supported by backend services in Python and FastAPI. Develop scalable integrations between manufacturing systems, robotics platforms, AMRs, sensor networks, and enterprise applications to enable intelligent factory operations. Design and implement digital twin environments using platforms such as NVIDIA Omniverse, Gazebo, or Unity Robotics Hub to support simulation, validation, and o
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA's Silicon Co-Design Group (SCG) leads the full product development lifecycle, from early architecture definition through silicon bringup to product release. The ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. This is where ideas become chips, and chips become products that define the state of the art — and we're building that future with some of the most motivated engineers in the industry. We're looking for a Senior Memory Systems Engineer to own HBM and LPDDR integration in sophisticated SoCs. This role covers the full stack, including silicon, package, embedded software, testing, and product development. The engineer will resolve the toughest system-level memory challenges throughout the process, building solutions that hold up at scale. What you'll be doing: HBM & LPDDR System Integration and Bringup: Drive HBM and LPDDR system integration, bringup, characterization, and debug for next-generation SoCs — taking memory subsystems through the full arc from first silicon to production-ready at scale. Full-Stack Memory Closure: <s
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
Electronic System Design and Analysis Engineer (Experienced or Senior) Company: The Boeing Company The Boeing Test Operations and Engineering (TO&E) team in Berkeley , Missouri is seeking an Electronic System Design and Analysis Engineer to join our Instrumentation Installation Design group to support all major programs in the St. Louis area. We are currently hiring for a broad range of experience levels including Experienced and Senior Electronic System Design and Analysis Engineers. Position Responsibilities: Electrical Design & Testing AC and DC power applications Assisting shop with build information Reading electrical drawings/schematics Troubleshooting Electrical component/wiring faults Production Functional and Environmental Acceptance Testing Designing new test sets for various weapon and aircraft electronics Performs system tests to verify operational and functional requirements Supports resolution of product integration issues and production anomalies Researches basic technologies for potential application to company business needs Assists in monitoring suppliers' performance to ensure compliance with requirements Develops or modifies basic hardware and software designs based on defined requirements Assists with the development and documentation of electronic and electrical system requirements This position is expected to be 100% onsite. The selected candidate will be required to work onsite at one of the listed location options. Basic Qualifications (Required Skill/Experience): <
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
As one of the technology industry's most desirable employers, NVIDIA has been redefining accelerated computing, computer graphics and leading the Artificial Intelligence revolution. NVIDIA's innovation is fueled by its great technology—and amazing people. We are seeking a Senior Silicon and System Product Lead to influence, innovate and take our next generation products to the market. As part of the Silicon Solutions Team, we architect and deliver groundbreaking system solutions that integrate all aspects of the system from silicon design, software design to operations and final deployment in multiple market segments that NVIDIA serves. This position offers an unique opportunity to collaborate with multiple organizations in the company and grow your career in a high impact role. We need a passionate, hard-working and creative individual to lead the products all the way from market analysis to delivering the features on the final product. What you'll be doing: Drive product performance and power targets, trade-off features/configurations and provide innovative solutions to complex silicon and system level problems. Evaluate new market segments and use cases; translate market requirements to engineering problem statements and metrics. Innovate Performance, power, yield and quality optimizations and features for the world’s fastest power-shipping products in the GPU and SoC market segments spanning gaming, automotive, datacenter and DL/AI. Develop methodologies and requirements for multi-functional teams to drive silicon and system product features to production. Incorporate productization feedback to improve the next generation. Lead the team for feature requirements and schedule from architecture to silicon phase of projects. Work alongside system architects, designers, marketing teams, chip and board designers, software/firmware engineers, HW/S
Senior System Software Engineer - Halos Core and Robotics Platform — US, CA, Santa Clara. Apply via Workday.
Senior System Software Engineer - Halos Core and Robotics Platform — US, CA, Santa Clara. Apply via Workday.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated
NVIDIA is now looking for a Senior Memory System Engineer to join our ASIC Memory Subsystem team! As a Senior Systems Engineer at NVIDIA, you'll join a group of hardworking engineers to develop and architect innovative Memory Solution for Tegra SoCs. In this position, you'll make a real impact in a multifaceted, technology-focused company. You will work with memory controller/PHY and Platform / System architect, Firmware, SI/PI, Memory suppliers to design and architect cutting edge, high speed and lower power memory technology for NVIDIA CPUs and SOCs. What You Will Be Doing: Analyze future DDR/LPDDR/HBM technologies to determine optimum performance, power, function and RAS in memory for Next generation SOC and Systems. Collaborate with ASIC Architects, Designers, Software and Firmware SW/FW teams to drive memory technology and associated requirements for memory controllers. Define Memory module, Package, and PCB layouts appropriate to the system workloads Debug and bring up memory evaluation / validation and failure issues on memory technology. Collaborate with DRAM suppliers and industry partners on to develop memory and memory related component technology. What We need to see: Bachelor's degree or master’s degree in Electrical Engineering, Computer Engineering (CE), or a related field (or equivalent experience) 10 years of proven track record in DRAM design, module design, or memory sub system design. Deep understanding and strong fundamental of memory design, features, ECC algorithm, SI and PI (Training algorithm) in DDR, LPDDR, and HBM. Strong understanding of memory sub system level interaction with Cache, Memory controller and PHY. Experience in the design, bring-up and validation for memory failure analysis Experience with Python, C/C
NVIDIA is a leading artificial intelligence computing company, and we are paving the way with innovations in self-driving cars, machine learning, supercomputing, gaming, and visualization. We give automakers, tier-1 suppliers, automotive research institutions, and start-ups the power and flexibility to develop and deploy breakthrough artificial intelligence systems for self-driving vehicles. Our unified computing architecture enables training deep neural networks in the data center, and then seamlessly runs them on NVIDIA DRIVE Platforms inside the vehicle. The Hypervisor and RTOS Team within NVIDIA DRIVE Software plays a critical role in NVIDIA's expansion into the world of artificial intelligence and autonomous vehicles. Our job is to facilitate the sharing and separation of system resources while achieving real-time, safety, and security requirements. We develop Hypervisor and RTOS with a strong focus on automotive quality, safety and security needed for the real-time, highly available system level components of world-class Autonomous Vehicles. We are making extensive use of formal methods to automate our workflow and increase the quality of our SW. We are hiring now for the position of Senior System Software Engineer for Hypervisor and RTOS What you’ll be doing: Design and develop new features for RTOS and hypervisor software stack. Bring up and optimize RTOS and hypervisor stacks on new NVIDIA Tegra SoCs. Develop high-integrity software using best-in-class engineering, safety, and security practices. Debug complex system-level issues across hardware, firmware, RTOS, and virtualization layers. Lead team-wide technical initiatives by building alignment, coordinating execution, and driving them to completion. What we need to see: BS, MS in CS/CE/EE or a related engineering field or equivalent experience </
Other cities to consider
More places hiring for this role
Get new senior system and manufacturing codesign architect jobs in United States by email
Daily job updates · Unsubscribe anytime