About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. The team ensures that large, highly complex infrastructure programs execute predictably across planning, design, construction, hardware deployment, commissioning, and operational handoff. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Supply Chain, and our external infrastructure partners to keep programs aligned, risks visible, and execution moving at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive execution across large-scale AI infrastructure deployments. This role is responsible for orchestrating cross-functional delivery programs spanning multiple organizations, ensuring dependencies remain synchronized from early planning through production readiness. You will develop operational mechanisms that allow Industrial Compute to scale infrastructure delivery across multiple campuses simultaneously. Rather than owning any individual engineering discipline, you will own program health—bringing together engineering, construction, operations, supply chain, and partner organizations into a single coordinated execution model. Success in this role requires exceptional program management, executive communication, systems thinking, and operational rigor. You should be comfortable operating amid ambiguity while bringing structure to highly technical, multi-year infrastructure programs. Candidates should have experience leading complex infrastructure, cloud, data center, semiconductor, networking, manuf
Jobiba hiring network
Data Center Ssd Performance Validation Engineer Jobs
8,341 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current data center ssd performance validation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
About the Role Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products. As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs. This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes — you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships. Responsibilities Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps
Job Details: Job Description: Intel is shaping the future of technology to help create a better future for the entire world. Our work in pushing forward fields like AI, analytics, and cloud-to-edge technology is at the heart of countless innovations. With a career at Intel, you'll have the opportunity to use technology to power major breakthroughs and create enhancements that improve our everyday quality of life. Join us and help make the future more wonderful for everyone. Want to learn more? Visit our YouTube Channel or the link below. Life at Intel The Role and Impact As a SoC Power Thermal Performance Validation and Optimization Engineer, you will play an essential role in ensuring Intel's products achieve optimal power, thermal, and performance benchmarks at the system-on-chip (SoC) level. In this position, you will develop and execute validation methodologies, optimize hardware/software solutions, and perform SoC-level debugging to address power, thermal, and performance challenges. Your contribution will directly impact Intel's ability to deliver competitive, high-performing products to market. Business Group This role is part of Intel's Data Center Group (DCG), which focuses on designing innovative solutions to power the next generation of computing platforms for data centers. As a member of the PTP PreSi Correlation Solution team, you will contribute to shaping the development and validation processes for cutting-edge technologies, ensuring Intel meets and exceeds industry standards for performance and efficiency. Key Responsibilities: Develop and execute SoC-level power, thermal, and performance validation and
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As part of Micron's Technology Engineering & Innovation (TE&I) organization, you will have the opportunity to shape the future of our global infrastructure platforms while enabling business growth, operational resilience, and digital transformation at scale. The Opportunity Micron is seeking a transformational Senior Director of Technology Engineering & Innovation (TE&I) to lead the strategy, engineering, operations, and modernization of our global infrastructure ecosystem. This role is responsible for defining and executing Micron's vision across enterprise networks, cloud platforms, data center strategy, database services, infrastructure engineering, automation, observability, and global infrastructure operations. As a key member of the TE&I leadership team, you will drive innovation, operational excellence, and strategic transformation while building a high-performing organization focused on business outcomes and exceptional customer experiences. The successful candidate will be equally comfortable developing multi-year technology strategies, leading large-scale infrastructure transformations, driving operational performance, developing talent, and fostering a culture of collaboration, accountability, and continuous improvement. What You Will Do Lead Global Infrastructure Strategy & Transformation Define and execute Micron's global infrastructure vision and strategy. Develop multi-year roadmaps for: Enterprise Network Services <l
The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. Roles and Responsibilities: Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, we are seeking a Senior Account Manager who will bring technical and business insight to help grow the Networking business with Cloud Service Providers, leveraging NVIDIA’s fast-growing Datacenter Networking Business. What you'll be doing: You will be an integral part of a small, forward-thinking team responsible for a significant and rapidly growing revenue contribution to the Enterprise products group, the fastest growing and most dynamic segment at NVIDIA. Working with leading CSP to grow deployment of NVIDIA networking solutions Define and drive the discussions, strategy and tactics for achieving revenue growth Partner and collaborate with technical teams on roadmap updates, new designs, competition Representing the customer’s strategy & needs to internal stakeholders and vice versa Delivering a concise state of the business at any given time What we need to see: Bachelor’s degree or higher in related field (or equivalent experience). 12+ years of Technical Sales or Product Management with focus on Data Center Networking for CSPs An understanding of the business and technology landscape of CSP Data Center Networking <l
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Manager, Electrical Engineering is responsible for the electrical systems of the KARNO generator, including high-voltage power electronics, battery systems, low- and high-voltage architecture, wiring harnesses, and the hardware that converts linear motion into electrical output. This is a working manager role: the position leads and develops a team of electrical engineers while remaining directly involved in technical execution, including circuit architecture, schematic review, and hardware bring-up in the lab. The Manager is accountable for the technical excellence, safety, and reliability of the electrical engineering function, and for establishing the design standards and review practices the team works to. The position plans team capacity, owns hiring and development for the electrical engineering staff, and partners with mechanical, controls, supply chain, and program management on system integration. KARNO systems are deployed in data center, military, and industrial applications. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Own critical electrical designs personally, including regular time at the bench and in the test cell, while leading the team as a practicing engineer. Lead the electrical engineering team in the design and development of KARNO generator electrical systems, including high-voltage power electronics, battery systems, linear generator power stages, and low-voltage controls hardwar
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems. You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory. You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership. You're good at bringing people together, mechanical, electrical, thermal, softw
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent builds AI and RISC-V compute hardware and licenses its technology to partners worldwide. The Corporate Development Associate leads client-facing analysis for the strategic deals side of the team, owning the evaluation, structuring, and execution of partnerships and commercial transactions. The role suits someone who takes on responsibility quickly, operates independently, and is credible in front of partners, customers, and the C-suite. Corporate Development is responsible for the transactions that shape Tenstorrent as a company: raising capital, acquiring or partnering with other businesses, and structuring the commercial and licensing agreements that extend the reach of its technology. The team runs Tenstorrent's equity and debt financings and its investor relations, evaluates and executes buy-side and sell-side M&A, negotiates strategic partnerships and joint development agreements, supports licensing and strategic sales, and develops government and sovereign AI programs. Counterparties include semiconductor and systems companies, data center operators, institutional and strategic investors, and government bodies in North America, Asia, and Europe. The team reports directly to Tenstorrent's Chief Strategy Officer and works day to day with the CEO and the executive team. It is small, and every member carries real responsibility
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent builds AI and RISC-V compute hardware and licenses its technology to partners worldwide. The Corporate Development Manager leads workstreams across the full range of the team's work: fundraising and investor relations, mergers and acquisitions, strategic partnerships, licensing, and commercial transactions. Managers own the evaluation, structuring, negotiation, and execution of transactions, lead the analysis behind them, and represent Tenstorrent directly with partners, customers, investors, and the executive team. The role suits someone who takes on responsibility quickly, operates independently, manages and develops junior team members, and is credible in front of the C-suite. Corporate Development is responsible for the transactions that shape Tenstorrent as a company: raising capital, acquiring or partnering with other businesses, and structuring the commercial and licensing agreements that extend the reach of its technology. The team runs Tenstorrent's equity and debt financings and its investor relations, evaluates and executes buy-side and sell-side M&A, negotiates strategic partnerships and joint development agreements, supports licensing and strategic sales, and develops government and sovereign AI programs. Counterparties include semiconductor and systems companies, data center operators, institutional and strategic inv
Get new data center ssd performance validation engineer jobs by email
Daily job updates · Unsubscribe anytime