Jobs in United States

Cpu Verification Fellow in United States

88 active opportunities · Updated October 2026

Explore current cpu verification fellow jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de

AWSRestAIRust
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our Tensix Team is building the future of AI compute with a ground-up architecture centered on scalable RISC-V processors. As we push performance boundaries, we’re reimagining the frontend of our RISC-V cores to deliver major gains in programmability, efficiency, and developer experience. This is a rare opportunity to shape the CPU architecture at the heart of our AI platform and lead one of the most strategic technical efforts at Tenstorrent. This role is hybrid, based out of Toronto, ON, Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced Microarchitect: 10+ years of deep expertise in CPU performance modeling and microarchitecture design. AI Workload Expert: Deeply familiar with the computational and memory bottlenecks of modern AI workloads, particularly Large Language Models (LLMs). Hardware-Software Co-Designer: Driven to architect custom instruction set extensions and validate their performance gains against real-world workloads. Ways to stand-out: Familiarity with open-source RISC-V cores, AI-based agentic workflow experience What We Need Profile & Analyze: Dissect cutting-edge AI workloads to identi

AWSAISEMHR
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

We are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute rack validation team. As a diligent leader in Systems Engineering, you will drive multiple aspects of validation throughout the life cycle of the program. In this high visibility position, you will be part of a leading team to innovate and improve system bring-up and enablement abilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The technical leader will be driving keys areas of system validation including leading first silicon & system bring-up (nodes and rack level systems) - rack level systems and blades will be based of ARM server architecture. Candidate will be immersed in challenging system enablement work, system validation (end-to-end) methodology, tests development and execution as well as triage/debug of critical issues to meet critical program milestones at POR quality. The candidate will also be a key contributor to state-of-the-art HW and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: Lead the systemenablement (including first silicon and other FW components) to ensure system capabilities are brought up as per plan of record and system architecture spec. Drive organization wide methodology for Firmware integration and best known configuration (HW/FW/SW) usage model by leading the release of deployment ready solutions. Develop key methodologies, lab HW and system SW capabilit

LinuxAIGoExcel
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to join our CPU design team to define and implement RTL for high-performance CPUs. You’ll work on a CPU based on RISC-V ISA, collaborating with DV, PD, and performance teams to deliver a functional, timing, and power-converged design. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced in CPU microarchitecture with expertise in Rename, Scheduler, ROB, Load Store, Branch Prediction, Cache or Datapath. Skilled in RTL coding (Verilog/VHDL) and familiar with industry-standard tools for simulation, synthesis, and power analysis. Proficient in debugging RTL/logic across multiple design hierarchies and pre/post-silicon environments. Background in microarchitecture definition, design specification, and performance-driven trade-off analysis. What We Need Own RTL design and microarchitecture development for a portion of a CPU block of a high-performance RISC-V CPU. Collaborate closely with DV, PD, and performance engineers to meet functional, timing, and power goals. Use innovative techniques to optimize power, performance, and

AWSAIGoSEM
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building next-generation processors and systems, bringing together world-class expertise across silicon, systems, and software. We are looking for a CPU Performance Modeling Architect to help evaluate, shape, and optimize the performance of future CPU architectures. In this role, you’ll use performance modeling, workload analysis, and deep understanding of CPU architecture to answer complex questions about how a processor should be designed. You’ll work closely with CPU architects, RTL designers, software and compiler teams, and system engineers to identify performance opportunities, evaluate architectural tradeoffs, and turn modeling insights into actionable design decisions. This role is hybrid, based out of Santa Clara, CA or Austin, TX. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A CPU architect, performance architect, or performance modeling engineer with experience influencing CPU architecture or microarchitecture decisions. You have a strong understanding of modern processor architecture and enjoy digging into why a CPU performs the way it does. You are comfortable combining hardware architecture, software, data, and

PythonAWSAIC++
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking an experienced Field Application Engineer to champion our revolutionary RISC-V CPU and AI accelerator IP products with customers worldwide. You will own end-to-end technical engagements, translating complex architectural advantages into customer wins while shaping our product roadmap through direct market feedback. If you combine deep technical expertise with exceptional customer engagement skills and want to drive the adoption of next-generation compute IP, join our team. This role is hybrid, based out of the United States or Canada near one of our main office locations. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A customer-focused engineer with proven success in CPU/GPU/NPU IP sales engagements. Technical expert who can articulate complex architectural benefits and conduct deep-dive analyses. Self-driven professional who independently manages multiple customer relationships from discovery to win. Strategic thinker who identifies customer pain points and positions solutions effectively. What We Need 7+ years in the semiconductor industry with a degree in EE or related field including customer-facing roles i

AWSAISEMHR
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s Silicon Co-Design Group is the team that gets every GPU, SoC, and CPU silicon program from first power-on to high-volume production. We are hiring a Senior Manager to lead our Test, Manufacturability, Reliability & Quality (TMRQ) organization. This is not a coordination role . Your work decides if a product can be built at scale and trusted in the field. These include production test development (SLT, BLT), control run flow, system reliability stress (HTOL), platform- and board-level manufacturing issue closure, and field diagnostic test development. You lead a team of individual contributors and a first-line manager at the layer where silicon, platform, and software collide with manufacturing reality. Decisions you make show up in yield curves, production ramp , and customer escapes. You are the leader the program turns to when a build is stuck, a control run is fallout-heavy, or a field return points back at silicon . The exceptional hire also uses AI deliberately — with demonstrated workflow impact and the judgment to know where it compresses real work and where it introduces risk. What you will be doing: Keep programs moving. Own the technical execution and velocity of SLT, BLT, Board/Chip/Rack CR, and system reliability stress (HTOL) across every GPU, SoC, and CPU silicon program. Close the hardest multi-functional failures. Resolve Vmin and binning escapes, performance shortfalls, and power anomalies by driving root-cause across design, methodology, DFT, ATE, package, software/firmware, and manufacturing — and own the WARs and productized fixes through to confirmation. Give leadership the clarity to act. Convert raw integration signals — CR fallout, BLT/SLT yield, SHTOL/CHTOL data, RMA trends, customer escalations — into decision-ready options that enable executive leadership to act with confidence on POR, QS/PS gates, a

T
📍 Austin, Texas, United States
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our diverse team of technologists have developed a high performance RISC-V-based CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance teams. This is the plumbing that both teams' performance work stands on — benchmarking automation, data collection, and workload capture across silicon, FPGA/emulation platforms and performance models. You'll work across both teams, enabling performance and software engineers to spend their time on analysis and optimization instead of running experiments by hand. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Great at identifying problems and developing solutions, with a bias toward owning the systems you build. Enjoys building tools and optimizing workflows so other engineers can move faster. Strong Linux systems engineer, c

PythonLinuxAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

Build the infrastructure that keeps every NVIDIA chip aligned from first spec to final shipment. NVIDIA's Silicon Co-Design Group sits at the convergence of architecture, silicon, systems, and manufacturing. The System–Manufacturing Architecture (SMAC) team coordinates between system specifications and manufacturing test specifications from pre-silicon POR through production release across GPU, SoC, and CPU programs. When that alignment drifts, silicon faces the consequences: escapes, yield loss, and performance loss. We're hiring a Senior Manufacturing & System Co-Design Workflow Engineer to lead the methodology and infrastructure that maintains holistic, systematic alignment, at scale across the full portfolio. The strongest candidates in this role design the workflow before being asked to fix a program, and build the checks and automation that confirm alignment holds long after they've moved on to the next problem. What you’ll be doing: SMAC Workflow Methodology: Define manufacturing spec types, including schema and semantics, derived from system PORs and features. Own the methodology that governs how specification work gets structured, versioned, and validated across the program lifecycle. Production Python Pipelines & Automated Checks: Develop production-grade Python pipelines and automated checks that catch specification drift between system POR and manufacturing test programs ,ATE, SLT, BLT, L10+, before silicon exposes the discrepancy. The goal is that misalignments surface in the workflow, not on the tester. E2E Program Integration & TPM Attestation: Wire SMAC work into the end-to-end program spine, milestones, gates, and artifacts, and define explicit TPM-driven attestation when checks lag. Alignment can't be assumed; it must be proven at every stage. Agent-Ready Tooling & CI Infrastructure: Integrate tooling into an agent-ready

N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA's accelerated computing platforms move data at speeds that push the limits of what silicon and physics allow. Whether a high-speed interface trains reliably, maintains accurate margins, and survives every platform topology it will ever see is a question we answer ourselves. This role does that work. The Silicon Co-Design Group leads the boundary between what was designed and what was built. When a GPU, CPU, or SoC ships with interfaces that work at scale, this team is the reason. Most engineers debug within a layer. You will own the full stack. When an interface fails to train, the link margin is unexpectedly tight, or a customer reports a critical silicon issue, you trace the problem through protocol behavior, signal integrity, firmware, platform topology, and silicon marginalities. You then confirm that the fix works. Your methodology shapes how NVIDIA validates high-speed interfaces across generations. Your decisions affect yield, production ramp, and field quality. This is not a coordination role. The engineers who do it well hold protocol depth and system breadth simultaneously, never lose the thread across hardware, firmware, and software, and have the judgment to know when to go deeper and when to act. They are rare. Should that describe you, read on. What you'll be doing: Own post-silicon bring-up, characterization, validation, and debug of PCIe, NVLink, C2C, and other HSIO interfaces across NVIDIA GPUs, CPUs, and SoCs from first power-on through production readiness. Close the hardest failures. Drive root cause across protocol behavior, signal integrity, firmware and driver interactions, platform topology, and silicon marginalities and own every fix through to confirmation. Define validation strategy. Set test coverage, debug priorities, margining methodology, and stress criteria f

T
📍 Austin, Texas, United States
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is looking for a mid- to senior-level Physical Design Engineer who will contribute to the physical design of high-performance chips for industry-leading AI/ML architectures, spanning implementation from synthesis through tapeout. You will partner with front-end and physical design engineers to optimize floorplanning, timing, power, performance, and area across multiple IPs. Along the way, you will build end-to-end ASIC expertise while learning from experienced engineers across the chip development process. This role is hybrid , based out of Austin, TX, Fort Collins, CO, or Santa Clara, CA . We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are An engineer excited to work on high-performance designs for industry-leading AI/ML architectures. A collaborative problem solver who enjoys working with experienced engineers across ASIC disciplines. Grounded in logic design fundamentals and gate- and transistor-level implementation. Curious about how early architectural and RTL decisions shape physical implementation and final chip quality. What We Need A BS, MS, or PhD in Electrical Engineering, Computer Engineering, Computer Science, or a relat

T
📍 Austin, Texas, United States
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is looking for a mid- to senior-level Physical Design Engineer with a strong background in static timing analysis (STA). In this role, you will help converge multi-million-gate designs on advanced process nodes and partner with cross-functional teams through final signoff and tapeout. This role is hybrid , based out of Austin, TX, Fort Collins, CO, or Santa Clara, CA . We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A physical design engineer with 8+ years of experience in STA, timing closure, and advanced-node design implementation. Experienced with timing flows, constraints, clocking, CDC, RDC, derates, uncertainties, guardbanding, and final signoff. Skilled in writing and debugging SDC constraints, creating ECOs, developing closure strategies, and analyzing worst-case corners. A strong collaborator and communicator who can drive results across distributed, cross-functional engineering teams. What We Need Drive timing convergence across blocks, partitions, sections, and top-level designs; generate and validate timing models and publish results. Analyze timing violations, develop and propagate fixes, debug timing-flow issues, and

I
📍 California, Santa Clara, United States
✓ High-confidence listingCompany trend +315.4%
Quick readStrong listing-quality and freshness signals

Job Details: Job Description: As a Senior Power and Performance (PnP) Engineer, you will be responsible for PnP product execution focused on measurements, analysis, and projections/estimates of Intel's unlaunched notebook and desktop products. Based on your technical knowledge of both Intel and the competition, you will help influence OEM customers, internal marketing teams, internal engineering teams, debug Intel silicon/platform on PnP issues, and be an integral part for Intel product launches. Core responsibilities to include following: Measure, analyze, and debug workloads to call out any power and/or performance gaps and close with internal engineering teams/architects Measure and analyze SoC/CPU and platform power on a rail by rail basis and debug any issues/gaps Align with other internal engineering teams and architects to ensure there is consensus on power and performance projections/estimates/measurements for both internal and external communication Guide, educate, and influence internal engineering teams, field account teams, and marketing teams on power and performance positioning of Intel products Create power and performance estimates on upcoming Intel products based on workloads to aid marketing decisions and help set internal KPI targets Own the publication of the Power Performance Guide (PPG) collateral to help guide OEM customers Bring up full test platform to enable both power and performance measurements including setup, calibration, and measurements. Qualifications: You mus

Recruitment
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

C$100K – C$500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking an Physical Design Engineer to lead cross-functional efforts to solve complex physical design challenges and develop end-to-end RTL-to-GDS methodologies across advanced nodes, with a strong focus on PPA and runtime improvements. The engineer will architect, integrate, and deploy AI/ML-driven solutions into production physical design flows, creating custom CAD tools and partnering with internal teams and EDA vendors to drive next-generation, ML-enabled capabilities. This role is hybrid, based out of Santa Clara, CA or Austin, TX or Fort Collins, CO. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who you are BS in Electrical or Computer Engineering (or equivalent experience) with 5+ years in Physical Design CAD methodology at advanced nodes. Proven track record improving PPA and/or runtime on high-performance, low-power taped-out designs. Hands-on with industry-standard EDA tools (e.g., Fusion Compiler) across synthesis, P&R, STA, signoff, and hierarchical flows. Strong Python/Tcl and data skills, with interest or experience in ML frameworks (PyTorch, TensorFlow), and the ability to drive complex projects independent

PythonAWSRestAI
🔔

Get new cpu verification fellow jobs in United States by email

Daily job updates · Unsubscribe anytime