About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Infrastructure Quality team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, general contractors, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. Our work spans from vendor qualification through commissioning, ensuring operational readiness across our global portfolio. About the Role We are seeking an experienced Manufacturing Quality Engineer (MQE) to establish, implement, and manage a manufacturing-focused quality program for datacenter infrastructure. This role will be responsible for vendor oversight, quality assurance, process improvement, and issue resolution for all critical systems. You will lead vendor audits, monitor performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced risks, and operational reliability. By partnering with vendors, construction teams, and internal stakeholders, you will help ensure OpenAI’s datacenters are delivered on time and built to the highest operational standards. Travel Domestic and international travel as needed (estimated 40–60%) to manufacturing sites, datacenter locations, and partner facilities. Key Responsibilities Vendor Oversight & Performance Management Conduct manufacturing evaluation, audits, and improve vendor performance across production, inspection, testing, and delivery phases. Develop and track quality metrics to assess manufacturing performance and identify trends. Partner with vendors to refine processes, training, and quality controls to mitigate risks before shipment. Program Development & Execution Develop and maintain a datacenter-focused m
Jobiba hiring network
Data Center Ssd Performance Validation Engineer Jobs
8,157 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current data center ssd performance validation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Senior Software Development Engineer in Test - Datacenter Server OS — CA, Santa Clara. Apply via Workday.
NVIDIA is seeking a Senior Technical Program Manager to join the CSP Engagements team, focused on deep technical engagement with hyperscale cloud service providers for NVIDIA’s next‑generation datacenter systems such as Vera Rubin NVL72. This role is intended for experienced systems and embedded software leaders—including software engineering managers, technical leads, or senior architects—who have led datacenter server and platform software programs and can operate as a trusted technical partner to hyperscale CSP engineering teams. As a member of the CSP Engagements team, you will act as the primary technical engagement leader between NVIDIA’s system software organizations and CSP platform, system software, and AI teams, ensuring alignment, readiness, and successful large‑scale deployment of NVIDIA‑based datacenter solutions. What you will be doing: Lead deep technical engagements with hyperscale CSPs as the primary NVIDIA point of contact for system software, firmware, and platform readiness for NVIDIA datacenter products. Partner directly with CSP system software, firmware, and infrastructure engineering leaders to align on software architecture, bring‑up plans, deployment readiness, and production requirements for NVIDIA‑based server and rack‑scale platforms. Represent CSP technical priorities internally, advocating for customer requirements and tradeoffs across NVIDIA’s system software, firmware, hardware, silicon, and product teams are aligned to customer needs, timelines, and constraints. Own the end‑to‑end CSP engagement lifecycle, from early technical alignment and pre‑production readiness through large‑scale deployment, escalation management, and sustained production support. Drive bi‑directional technical communication: translating CSP system‑level requirements into actionable focus areas for NVIDIA engineering teams, while clearly communicating N
NVIDIA is seeking a System Test Engineer to join our Test Solutions Group, where you will develop and deploy manufacturing test solutions for next-generation datacenter systems. You will combine hardware and software expertise to expand test coverage, automate diagnostics, and solve complex production challenges at scale. NVIDIA has been reinventing computing for more than two decades. From pioneering the GPU and transforming computer graphics and gaming to powering the AI revolution, we continue to tackle problems that matter. Join a team dedicated to amplifying human creativity and intelligence. What you’ll be doing Define, develop, and implement manufacturing test solutions for datacenter system products. Investigate and introduce new test technologies and methodologies to improve quality, coverage, throughput, and production efficiency. Develop test strategies for new product features, including test identification, specification, fixture and equipment design, diagnostic software development, and qualification. Provide Design-for-Test (DFT) feedback early in the product design cycle. Debug complex hardware, firmware, and software issues throughout prototype and production build phases. Build and improve automated diagnostic-generation and validation infrastructure. Collaborate across engineering, operations, and manufacturing teams to deliver robust, scalable test solutions. Coordinate technical work across team members and drive execution on shared deliverables. What we need to see Bachelor’s degree in Electrical Engineering or equivalent practical experience. 3+ years of post-silicon validation experience with board-level or system-level products. Experience developing or supporting platform-level manufacturing test programs. </l
The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. NVIDIA’s Silicon Co-Design Group is hiring a Chip Lead to serve as the technical lead for one of our most consequential silicon programs! This is not project management, and it is not a senior IC role. You are the person the program partners with on its hardest technical questions — when the right answer is not obvious, and when leadership needs a single technical perspective to align on direction. You are accountable for the technical integrity of the chip end-to-end. You guide co-design feature integration, help resolve the toughest multi-functional bugs, and serve as the project-specific custodian of the qualification playbook. The program runs cleaner because you are on it! What you’ll be doing: You will partner across design, validation, software, and manufacturing to keep the program’s technical narrative clear and on track. Day to day, you will: Serve as the single technical point of contact for multi-functional decisions, issues, and trade-offs. Co-lead program-level feature integration from chip to system, surfacing inter-function dependencies and guiding them to resolution. Help resolve the program’s hardest multi-functional bugs by translating ambiguous, multi-team symptoms into root-cause closure on areas such as HBM, power and thermal, high-speed I/O, and packaging. Steward the qualification playbook. When the playbook does not fit a situation, guide the mitigations and capture the lessons as reusable methodology for other SSG programs. Shape the program’s technical narrative by surfacing key risks, trade
The Silicon Co-Design Group (SCG) sits at the crossroads of architecture, design, marketing, operations, and productization. Our work spans early architecture through final product delivery across Datacenter, Gaming, Robotics, Automotive, and Embedded markets. We work closely across functions to deliver chips that change what is possible. System Integration sits at the intersection of all of them. It is the layer where every architecture, design, software, and manufacturing decision meets reality. When something breaks late in a program, it usually breaks here first. We are hiring a Senior Manager to lead this team in the US and partner closely with teams globally. Your work will sit on the critical path of every NVIDIA silicon program, and the bar you set for system integration is the bar we ship to! The two hardest, highest-leverage problems in this seat: Find critical silicon issues earlier — often before software is production-ready. Left-shifting post-silicon coverage is the highest-value thing System Integration can do. Standing up wide-area testing as a repeatable capability is a core part of the role. Keep programs on milestone when upstream dependencies slip. Validation plans collide with reality every program! The team needs new strategies, not just contingency plans, to keep moving when software, firmware, or methodology slip. You will design and run those strategies. What you’ll be doing: Plan and execute post-silicon feature integration, PVT validation, and wide-area testing across NVIDIA’s GPU, SoC, and CPU programs. Build wide-area and in-system test as a repeatable capability that shifts post-silicon coverage left, so issues surface before we are production-ready. Lead resolution of the most complex system-level issues, RMAs, and HW/SW interaction problems with creative workarounds and focused lab experimentation. Deve
SCG sits at the crossroads of design, architecture, marketing, and productization—owning the journey from the architecture stage through final product definition across Gaming, Datacenter, Automotive, and Embedded markets. As a System Verification CoDesign Engineer, you will work on system-level speed features, develop the verification collaterals and automation infrastructure to characterize and validate them, and lead debug of the complex silicon issues that stand between a program and on-time shipment. This is a hands-on role for an engineer who combines deep technical craft with the drive to compress cycle time using modern tooling—including AI—without losing rigor. What You’ll Be Doing: Collaborate cross-functionally with system architects, hardware, firmware/software, process/reliability, and operations teams to co-design system-level speed features and deliver industry-defining products. Understand system level behavior and speed reliability margins, bounding box constraints and identify solutions that optimize margins . Translate hardware features and architectural requirements into verification techniques that achieve full coverage across testing flows. Perform closed loop validation by correlat ing silicon behavior against timing simulation and design expectations; provide actionable feedback to improve future designs. Define, prototype, and refine pre- and post-silicon bring-up flows to ensure
Job Details: Job Description: This position is part of Intel Datacenter Group's Office of Strategic Intelligence and Analysis, which focuses on delivering high-quality insights and business intelligence solutions to support the company's strategic goals. The team works across functional areas, leveraging analytics and business intelligence to drive impactful decision-making and operational efficiency. Join a group that values collaboration, customer-centricity, and cutting-edge data approaches to shape Intel's success in a rapidly evolving market. Position Overview As a Business Intelligence Analyst, you will play a pivotal role in transforming complex business data into actionable insights that drive decision-making in product roadmap decisions and improve business practices. You will collaborate with diverse stakeholders to translate market trends into actionable business strategy, craft compelling data stories, and deliver strategic recommendations that influence senior leadership product decisions. Your expertise in data analysis and visualization will empower the organization to navigate dynamic industry landscapes and achieve meaningful business impact. Key Responsibilities • Gather, clean, and analyze business data from various sources to identify patterns and actionable insights in support of business product decisionmakers. • Translate complex datasets into compelling narratives and visualizations, including interactive dashboards, to enhance data storytelling. • Pro
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary As a research engineer at Graphcore, you will contribute to the advancement of AI research, investigating new ideas that push the limits on important AI/ML problems. Specialised hardware has been the key driver of the progress of AI over the last decade, and we believe that hardware-aware AI algorithms and AI-aware hardware developments will continue to be critical to advancing this exciting field. We are therefore looking for individuals who combine strong machine learning experience with practical engineering skills to deliver impactful AI research. We are seeking AI researchers with strong software engineering experience, particularly in lower-level programming and performance optimisation for hardware efficiency. Our research spans a broad range of topics, including efficient training and inference, world models, life sciences, reinforcement learning, and beyond. You will work closely with researchers to generate ideas and translate them into scalable implementations, contributing to publications and projects that help to steer the future of AI hardware. The Team Graphcore Research participates in both fundamental and applied research, to characterise the computational requirements of machine intelligence a
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will lead the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with cross-functional teams, you will contribute to the design and implementation of SLT hardware, software, automation, and characterization solutions that support silicon bring-up, yield learning, manufacturing readiness, and production deployment. The ideal candidate will possess strong technical depth in semiconductor test and validation, a passion for solving complex engineering challenges, and a strong focus on product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Lead development and deployment of SLT hardware and software solutions supporting silicon bring-up, charac
Manufacturing Test Engineering Manager Position Summary We are seeking an experienced Manufacturing Test Engineering Manager to lead the development and execution of the end-to-end manufacturing test strategy for next-generation AI server platforms and datacenter infrastructure. This role is responsible for defining and driving the manufacturing test architecture from L6 board assembly through L11 rack-level integration and final system validation , ensuring world-class product quality, manufacturability, and production scalability. This leader will manage a team of 3–5 Manufacturing Test Engineers while partnering closely with Hardware, Firmware, Platform, Validation, Quality, Operations, and Joint Design Manufacturing (JDM) partners. The role owns the manufacturing test strategy, test coverage, factory test infrastructure, manufacturing capacity planning, and continuous improvement of manufacturing quality. Key Responsibilities Manufacturing Test Strategy Define and own the end-to-end manufacturing test strategy from L6 board assembly through L11 rack integration . Develop standardized manufacturing test methodologies that optimize quality, throughput, cost of test, and scalability across multiple products and JDM sites. Establish manufacturing test standards, best practices, and engineering processes that support high-volume server manufacturing. Technical Leadership & People Management Lead, mentor, and develop a team of 3–5 Manufacturing Test Engineers supporting multiple hardware programs. Establish team priorities, allocate resources, and ensure successful execution of manufacturing test deliverables. Foster a culture of technical excellence, accountability, collaboration, and continuous improvement. Serve as the primary technical escalation point for manufacturing test and production issues. Cross-Functional Engineering Collaboration Partner with Hardware, Platform, Firmware, Validation, Reliability, Quality, and Operations teams to ensure manufac
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced Silicon Test Engineer to join our Product Test and Diagnosis Department (PTD). This is a pivotal role and will involve building a team of engineers to develop System Level Test (SLT) capability within the company. Working closely with a cross-functional team you will implement SLT tests for a family of next generation AI Processors. The ideal candidate should have a focus on quality and demonstrate a good understanding of the importance of production test on the success of a product. T hey will have a proven Functional Test or ATE Test Engineering background, and will have a pragmatic, hands-on and flexible approach to a fast-changing environment. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Managing a team of
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary As a Senior Machine Learning Engineer in the Applied AI team at Graphcore, you will contribute to advancing AI technology by developing and optimising AI models tailored to our specialised hardware. You will work on large scale systems where performance is critical to the success of our projects. Working closely with the Software development and Research teams, you will play a critical role in identifying opportunities to innovate and differentiate Graphcore’s technology. We seek engineers with strong technical skills and an understanding of AI model implementation at scale, eager to make a tangible impact in this rapidly evolving field. The Team The Applied AI team’s role is to be proxies for our customers, we need to understand the latest AI models, applications, and software to ensure that Graphcore’s technology works seamlessly with the AI ecosystem and at scale. We build reference applications, contribute to key software libraries e.g. optimising kernels for efficiency on our hardware, and collaborate with the Research team to develop and publish novel ideas in domains such as efficient compute, model scaling and distributed training and inference of AI models for multiple modalities and applications. If you're excited about advancing the next gen
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
Get new data center ssd performance validation engineer jobs by email
Daily job updates · Unsubscribe anytime