Jobiba hiring network

Senior Cpu Design Verification Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior cpu design verification engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

We are seeking a qualified Senior Software Tools Development Engineer to join our GPU SWQA team. The successful candidate will have strong experience applying AI technologies to automate test cases and a deep understanding of Windows operating systems. Extensive knowledge of GPU, CPU, SoC, x86, and ARM architectures is required, along with expertise in PC I/O architecture and common bus interfaces such as PCIe, USB, and SATA. Familiarity with specifications for general PC architecture components is a plus. What you’ll be doing: Design and implement automated tests incorporating AI technologies for NVIDIA's device driver software and SDKs on windows platforms. Build tools/utility/framework in Python, C# or equivalent which would help automate and optimize the testing workflows in GPU domain. Develop and carry out automated and manual tests, analyze results, identify and report defects. Rigorously drive test automation initiative. Build innovative ways to automate and expand our software testing. Expose defects and constraints; Isolate and debug the issue(s) and find the root cause; Contribute to the solution and drive to closure. Measure code coverage for the software under test, analyze and drive code coverage enhancements. Develop applications and tools that accelerate development and test workflows and write fast, effective, maintainable, reliable and well documented code. Generate and test compatibility across a range of products and interfaces and validate different key software applications across a test matrix designed to test both breadth and depth. Provide peer code reviews including feedback on performance, scalability and correctness. Report test coverage and Go/No-Go status for deliverables, escalate critical issues, and drive them to closure. Participate in root cause analysis and corrective actions to continuo

pythonsqlmongodb
View job →
N
11 days ago

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

pythonkuberneteslinux
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA's GeForce Now, the next-generation gaming service powered by NVIDIA GPUs in the cloud, transforms a Mac, any PC, or just a mobile device into a high-performance gaming rig. GeForce NOW automatically keeps games up-to-date, and users around the globe can instantly stream the latest games in high-definition resolution at the lowest latency for the smoothest gameplay. Just click and play! Visit us at https://www.nvidia.com/en-us/geforce-now. In addition to gaming, the state-of-the-art low-latency streaming technology has expanded to a new range of applications, including augmented and virtual reality, artificial intelligence, and robotics. We are now looking for a Senior Software Engineer – Streaming with strong C++ skills and a deep interest in streaming technologies to join a team of highly skilled and motivated software engineers who help build the next generation of applications with media and data streaming capabilities. Now, are you passionate about driving this technology to its edge? Do you understand various streaming protocols? Can you solve complicated problems and propose innovative solutions? Then, we are keen to hear from you. What you’ll be doing: Design, develop, optimize, and debug C++ software for ultra low-latency streaming systems. Build and improve media streaming pipelines using technologies such as WebRTC, GStreamer, RTP/RTCP, and related protocols. Analyze CPU, memory, networking, media pipeline, and scheduling behavior across real-world workloads. Work across Windows, Linux, QNX, applications frameworks, and embedded platforms. Define and implement KPIs for networking quality, streaming quality, and user experience. Integrate streaming softwa

linuxartificial intelligenceai
View job →
N
1mo ago

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which the GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As a Developer Technology Engineer you will be at the forefront of innovation, working with leading industry partners and exciting OSS projects to accelerate RTX & DGX SoC performance for agentic use. This role offers an outstanding opportunity to collaborate with world-class talent and make a significant contribution to the next era of enterprise and consumer AI. What you'll be doing: Own key engagements with our fast-paced developer ecosystem partners, optimizing their applications to deliver outstanding end-to-end performance on RTX and DGX SoCs. Take ownership of performance optimization for compute-intensive CPU, AI, and 3D graphics workloads, working with domain experts to identify bottlenecks and implement effective solutions. Work across system software, GPU driver, architecture, and NVIDIA Research teams to influence next-generation, high-performance SoC platforms by bringing real-world workflows and actionable insights from partner and customer needs. Provide technical guidance and mentorship to junior engineers while contributing to an inclusive and high-performing team environment. What we need to see: BS or MS degree in Computer Science, Engineering, Mathematics or related degree (or equivalent experience) 5+ years of respective work experience as software developer Proficiency in C/C++, Python, software

pythonlinuxai
View job →
N
1mo ago

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li

pythondockerkubernetes
View job →
S
Snowflake
📍 Bellevue• Full-time• Remote
1mo ago

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Capacity Engineering At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic, fast-moving environments and approach challenges with an experimental mindset, rapidly testing emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s infrastructure is expanding rapidly across AWS, Azure, and GCP. The Capacity team plays a pivotal role in provisioning the cloud resources essential for Snowflake's operations and ongoing growth. Capacity Engineering accurately models demand, forecasts requirements, and delivers optimal CPU and GPU capacity on schedule. We drive hardware cost-efficiency and price/performance while continually maximizing fleet utilization. To achieve this across all major cloud providers, the team is building a centralized, self-se

REMOTEpythonjavaaws
View job →
R
Roblox
📍 San Mateo• Full-time• From $196.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Game Engine team at Roblox works on the systems that power the experiences at the heart of the metaverse. Our code is the technical foundation in our client, editor, and our simulation servers. The Game Engine team is broadly split into departments — Audio-Video-Communication, Avatar, Core AI, Digital Matter, Productivity, and Systems. This position is for the Systems Runtime pod. The Runtime pod builds and maintains the foundational C++ components that power the entire Roblox engine and Studio stack. We own the runtime layer — the performance, efficiency, and usability of the core primitives that every other engineering team relies on. Think of Runtime as the “standard library” and execution engine for Roblox's C++ world. Our ownership centers on three pillars: Concurrency System — our task scheduler and fiber runtime power how work is parallelized across CPU cores, letting teams scale features across platforms while keeping code understandable and debuggable. Memory System — we own the engine's memory allocator stack, tracking pipelines, and observability tooling to make memory behavior predictable and prevent leaks and fragmentation. Profiling and Observability — we build and evolve

awsgitrest
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role We’re looking for an Autonomy Engineer focused on onboard autonomy—the software that runs on the robot/vehicle/embedded computer and makes real-time decisions using onboard sensors and compute. You’ll build and ship reliable autonomy features that operate under tight latency, compute, and safety constraints in the real world. What You’ll Do Develop, integrate, and deploy onboard autonomy behaviors (e.g., navigation, obstacle avoidance, lane/route following, docking, interaction behaviors). Implement and maintain real-time decision-making components: behavior planning, state machines/behavior trees, local planning, and control interfaces. Build robust sensor-driven autonomy pipelines on-device (camera, lidar, radar, IMU, wheel odometry, GNSS), including synchronization, calibration hooks, and fault handling. Optimize autonomy performance for latency, CPU/GPU usage, memory, and power on embedded compute (e.g., NVIDIA Jetson,

pythonlinuxai
View job →
N
Nuro
📍 Mountain View• Full-time• From $193.9K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors About the Role Nuro is seeking a Software Engineer with expertise in large-scale infrastructure, workload orchestration, and data processing to join our ML Infrastructure team . In this role, you will focus on building and evolving the core platform that provides researchers and engineers with seamless access to compute and data resources. You will be responsible for executing the technical strategy for automated resource provisioning, high-performance workload scheduling, and efficient feature management to accelerate the Nuro Driver™ development lifecycle. About the Work You will build the foundation that powers Nuro’s model development from experimentation to production. Key responsibilities include: Resource Provisioning & IaC: Scaling automated infrastructure-as-code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments. Intelligent Scheduling: Designing and optimizing workload orchestration to maximize

redisawsazure
View job →
R
Roblox
📍 San Mateo• Full-time• From $280.5K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. About the Role: AI models reshaping how our community creates, plays, and connects, all run on Compute Platform. As Senior Product Manager, Compute Platform , you'll set the strategy and roadmap for Roblox's next-generation AI infrastructure: the rapidly growing fleet of GPUs and AI accelerators spanning Roblox core and edge data centers, and public cloud that decides how fast we can train, serve, and scale every model on the platform. You'll own the products that turn raw GPU hosts into reliable, production-ready AI compute - driver and firmware management, fleet-wide health and performance, and the abstractions product teams across Roblox build on. You Will: Drive strategy and roadmap for Compute Platform spanning Managed Kubernetes (Roblox Kubernetes Service), Managed Compute Services and other critical distributed systems, and our fleet of GPU and CPU machines managed via unified Fleet APIs - all across on-prem and cloud. Drive the evolution of our Compute infrastructure to support Roblox’s most critical workloads - from AI to Storage to Data Analytics and more - each with their own distinct requirements. Build and scale our GPU infrastructure to support training and inferen

awsazuregcp
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s cloud spend is in billions of dollars per year. Hence, it is critical for us to govern and optimize our cloud spend, both for margins and long-term competitive advantage. Cloud Efficiency team’s charter is to build scalable products that enable governance, monitoring and optimization of cloud spend. Think of this as Observability for cloud costs and efficiency. The team’s vision is to “Transform cloud spend into a competitive advantage by empowering teams to continuously optimize the per-unit cost.” In order to improve the overall cloud efficiency (i.e. cost per unit), it is critical to build monitoring products that collate costs with other factors such as utilization, attribution, hardware performance and architecture. Hence, there is an opportunity to build a unified, self-serve cloud efficiency product across Snowflake, that delivers actionable, real-time efficiency datasets through streamlined user experiences. This will enable thousands of engineers at Snowflake and will elevate cloud efficiency at Snowflake for long-term success. When developing these solutions, we think about the problem end-to-end: how do we collect data from different stacks (e.g. costs from AWS, GCP, Azure and CPU, Memory, Utilization) across Snowflake reliably, how do we store it eff

pythonjavaaws
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

pythonkuberneteslinux
View job →
I
11 days ago

Job Details: Job Description: As a Senior Power and Performance (PnP) Engineer, you will be responsible for PnP product execution focused on measurements, analysis, and projections/estimates of Intel's unlaunched notebook and desktop products. Based on your technical knowledge of both Intel and the competition, you will help influence OEM customers, internal marketing teams, internal engineering teams, debug Intel silicon/platform on PnP issues, and be an integral part for Intel product launches. Core responsibilities to include following: Measure, analyze, and debug workloads to call out any power and/or performance gaps and close with internal engineering teams/architects Measure and analyze SoC/CPU and platform power on a rail by rail basis and debug any issues/gaps Align with other internal engineering teams and architects to ensure there is consensus on power and performance projections/estimates/measurements for both internal and external communication Guide, educate, and influence internal engineering teams, field account teams, and marketing teams on power and performance positioning of Intel products Create power and performance estimates on upcoming Intel products based on workloads to aid marketing decisions and help set internal KPI targets Own the publication of the Power Performance Guide (PPG) collateral to help guide OEM customers Bring up full test platform to enable both power and performance measurements including setup, calibration, and measurements. Qualifications: You mus

recruitment
View job →
T
Tenstorrent
📍 Australia• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a highly technical Senior Staff/Principal Engineer to lead the porting and enablement of critical AI workloads. You will be a primary driver in migrating compute workloads to RISC-V architectures, ensuring our hardware is optimized for real-world application performance. The ideal candidate has a strong background in DevOps, workload porting, or application enablement . While this is an individual contributor role at its core, you will have the opportunity to grow and lead a small, specialized team over time as our workload migration efforts scale. This role is remote, based in the United States or Australia. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A Technical Catalyst: You thrive on the challenge of driving AI hardware porting to RISC-V and making complex software stacks run efficiently on new hardware. Systems Expert: You possess deep knowledge of system software, compilers, or low-level OS internals. You are an expert in ARM or x86 environments and are ready to apply those skills to the RISC-V frontier. A Project Driver: You have the technical authority to lead the implementation of a compute migration to RISC-V through

awsaidevops
View job →
T
Tenstorrent
📍 Austin• Full-time• $100K – $500K/yr
15 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are seeking a Senior Cost Accounting Manager to join our world-class accounting team. You will play a critical role in owning manufacturing cost accounting — including inventory, fixed assets, and COGS — building out standard cost processes, and partnering cross-functionally with Supply Chain, Sales, and R&D to drive accurate, scalable financial reporting as we grow toward IPO readiness. This role is hybrid, based out of Austin, TX; Santa Clara, CA; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Accounting or Finance degree; CPA strongly preferred. 8+ years in manufacturing or hardware cost accounting. Deep U.S. GAAP knowledge for inventory, fixed assets, and COGS. Expert with standard costing, variance analysis, and cost roll-ups. Comfortable turning Supply Chain, Sales, and R&D activity into accurate results. What We Need Own inventory, fixed asset, and COGS accounting under U.S. GAAP. Build and scale the cost accounting function for IPO readiness. Maintain and improve standard costs, updates, and cost roll-ups. Analyze manufacturing variances and partner with Operations on actions. Lead COGS and mar

awsaisem
View job →
🔔

Get new senior cpu design verification engineer jobs by email

Daily job updates · Unsubscribe anytime