Jobiba hiring network

Data Center Ssd Performance Validation Engineer Jobs

8,341 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current data center ssd performance validation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

YG
Yondr Group
📍 London• Full-time
16 days ago

About Yondr Yondr is a disruptor. We challenge convention and simplify complexity. A global developer, owner operator and service provider of data centers, we deliver complex data center capacity needs for the world’s largest tech companies. Our exponential growth sees us looking for extraordinary people to help accelerate us towards our vision: a tomorrow without constraints. But we can’t do this without you. About the Role With continued growth across our business, we are actively seeking to have a new member join our team. This is an exciting role to join an existing group of skilled development and investment professionals, enhancing the team capabilities and becoming an integral part of a rapidly growing business. We are searching for someone who has excellent quantitative and analytical skills, is comfortable building financial models and has experience with communicating with internal and external stakeholders. You will be a highly valued and trusted member of our ambitious and hardworking team, in a business that continues to scale to new heights. Main Responsibilities / Take responsibility for development and management of financial models and valuation analysis in support of potential transactions and investment opportunities / Support in all aspects of financing processes included due diligence, project management, data room management, financial analysis, legal negotiations and preparation of materials / Work closely with the Development and Finance teams to perform detailed financial analysis on new developments as well as existing portfolio assets / Production of presentation material for senior management, board members, investors, lenders and joint venture partners / Assist with commercial negotiation of legal contracts for transactions / Research and analysis of market trends and sector / asset valuations / Support in business strategy and individual project delivery /

gitrestai
View job →
YG
16 days ago

About Yondr Yondr is a disruptor. We challenge convention and simplify complexity. A global developer, owner operator and service provider of data centers, we deliver complex data center capacity needs for the world’s largest tech companies. Our exponential growth sees us looking for extraordinary people to help accelerate us towards our vision: a tomorrow without constraints. But we can’t do this without you. The Role The Critical Facility Shift Technician is responsible for the safe, reliable, and efficient operation of mission-critical infrastructure within the data center. This role covers all aspects of energy isolation, switch activity, and facility rounds, ensuring compliance with safety standards and operational procedures. The Critical Facility Technician reports directly to the Critical Facility Shift Manager and works as part of a larger team that strives for excellence through teamwork and accountability. This is a shift-based role requiring flexibility to work nights, weekends, and holidays as part of a 24/7 operations team. Main responsibilities Operations & Rounds Conduct scheduled and unscheduled facility rounds, inspecting all critical systems (electrical, mechanical, HVAC, fire/life safety, and security) for proper operation and signs of abnormality. Monitor Building Management System (BMS), Electrical Power Monitoring System (EPMS), and other automation platforms for alarms and alerts; respond promptly per SOPs. Maintain accurate logs of alarms, operational events, isolations, and shift activities. Ensure all Statuory PPM workorders are completed ontime and any resulting actions are logged and tracked and remediated in good time . Ensure all Yondr self delivered works are covered by acurate and relevant risk and method statements. Confirm that all contractor RAMS are accurate befor

restaigo
View job →
YG
Yondr Group
📍 London• Full-time
16 days ago

About Yondr Yondr is a disruptor. We challenge convention and simplify complexity. A global developer, owner operator and service provider of data centers, we deliver complex data center capacity needs for the world’s largest tech companies. Our exponential growth sees us looking for extraordinary people to help accelerate us towards our vision: a tomorrow without constraints. But we can’t do this without you. About the Role A very exciting opportunity for a finance role within the business. The role will provide support in relation to end to end responsibility for operating companies (“OpCos”), including financial reporting and analysis across the business. Main Responsibilities / Perform all related accounting functions associated with OpCos, including account reconciliation and financial statements / Perform monthly variance analysis, identify trends, and make recommendations for improvements / Ensure the accounting policy applied is correct for all entities and consistent per the Group Reporting framework. / Manage various third parties, including auditors and outsourced services. / Interface and collaborate with other departments on monthly draws, cash management, tax filings, along with other items. / Supporting the Group OpCo Financial Controller with any activities as required. / Identify and drive process improvements, including the creation of standard and ad-hoc reports / Actively be visible to the group functions, acting as the go-to contact for the OpCos / Ensure OpCos adhere to the Group Operating Model, including compliance with transfer pricing policies and governance requirements / Act as a key business partner across Finance and the wider organisation to deliver accurate financial reporting, support decision-making and drive process improvements Qualifications and experience / Qualified Accountant (ACA / ACCA / CIMA / CA) / 3-4+ experience of working in a finance role / Strong presentation, analytical skills, p

awsrestai
View job →
G
16 days ago

Mechanical and Thermal Laboratory Technician Position Summary Graphcore is a globally recognized leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data center hardware that provide the specialized processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. We are opening a new AI Engineering Campus in Austin, which will play a central role in Graphcore's work building the future of AI computing. Responsibilities The Mechanical and Thermal Laboratory Technician is a hands-on technical role supporting the development and validation of advanced AI hardware systems for data center environments. Working as part of a cross-functional engineering team, this individual will be responsible for executing mechanical and thermal laboratory testing, supporting product validation activities, prototype fabrication and assisting with troubleshooting and root-cause analysis of complex hardware systems. Requirements Associate degree in Mechanical Engineering Technology or a related technical field preferred. Equivalent combinations of education, training, and relevant experience will be considered, including experienced non-degreed candidates or candidates with degrees in unrelated disciplines. Minimum of 5 years of experience working in mechanical laboratories, machine shops, test labs, or similar technical environments. Experience with server hardware platforms and data center equipment. Knowledge of Direct Liquid Cooling (DLC) systems and their implementation in server environments. Experience operating forklifts, pallet jacks, and other material-handling equipment. Key Responsibilities Execute mechanical and thermal test plans to validate hardware designs a

aisemtraining
View job →

By submitting your resume, you’re expressing interest in our 2027 RDSS (Research and Development Substitute Services) program. Please confirm your eligibility with the local district office before applying the role. NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU & SOC acts as the brains of computers, robots, and self-driving cars that can understand the world. We are looking for Design Validation Engineer in Taipei for board/system power qualification and function test, responsible for NVIDIA Data Center platform, Graphics board, ARM Based platform and Autonomous Driving Platform. If you're creative and autonomous, we want to hear from you! What you'll be doing: Server/Motherboard/Mobile system functionality test. Clock/Sequency/GPIO signal measurement and verification. WAT (wide area test) function test. Gaming test. High/Low speed interface SI test. Co-work with hardware design engineers on debug and FA. Co-work with mechanical/thermal engineers for cooler/heatsink design. What we need to see: Masters degree in EE, computer science, or relative majors. Proficient with Linux system and GPU setup. Capable of writing simple shell scripts and analyzing test results using commands. Familiar with PC and Datacenter system hardware assembly, MB BIOS upgrade, and network troubleshooting Proficient with a programming language (C/C++, Python, Java, or Perl) Understand UEFI boot OS startup process and out-of-band (

pythonjavalinux
View job →
N
Nvidia
📍 Santa Clara, United States
1mo ago

NVIDIA is seeking an a PCB Library Engineer to join our PCB Design Infrastructure team. In this role, you will help develop and maintain the PCB library assets used across NVIDIA's Data Center, AI, Networking, Automotive, and Graphics products. Working alongside experienced PCB designers, library engineers, mechanical engineers, manufacturing engineers, and component engineers, you will create and validate component footprints, schematic symbols, mechanical components, panel definitions, and other critical design assets that enable successful product development. This position provides an excellent opportunity to build expertise in PCB design, manufacturing, component engineering, and design automation while supporting some of the most advanced computing platforms in the world. What you'll be doing: Develop PCB footprints, padstacks, schematic symbols, and mechanical library content using Cadence PCB design tools. Review component datasheets, package drawings, and engineering specifications to create accurate design libraries. Support library verification, release, and documentation processes. Partner with PCB design, mechanical engineering, manufacturing engineering, and operations teams to resolve library-related issues. Learn and apply industry standards, including IPC requirements, DFM, DFA, and DFT principles. Support quality initiatives to ensure library content is accurate, manufacturable, and scalable. Participate in continuous improvement and automation efforts within the library environment. Develop a strong understanding of PCB fabrication, assembly, and component technologies. What We Need to See: BS degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Manufacturing Engineering, or a related field or equivalent experience. <p

pythonai
View job →
N
1mo ago

NVIDIA Networking is a leading provider of innovative end-to-end InfiniBand and Ethernet connectivity solutions for servers and storage. Our portfolio includes adapter cards, switches, cables, and software designed to optimize Data Center performance with industry-leading bandwidth and scalability. We serve diverse sectors such as high-performance computing, enterprise, cloud computing, and Web 2.0. Our mission is to stay ahead of the market by delivering groundbreaking products and services. Our Ethernet solutions are tailored for industries like Media & Entertainment and any domain that benefits from advanced DataStream and TCP/IP acceleration. What You’ll Be Doing: Lead a team of 8&#43; mechanical design engineers. Define priorities, create project plans, and allocate resources for mechanical programs in coordination with Product Managers. Drive all electro-mechanical, automated JIG and thermal design aspects, of production test setups, ensure readiness of test and assembly infrastructure for high-volume manufacturing. Develop multiple early design concepts in fast-paced product development cycles. Lead task forces to investigate and resolve production issues, reliability concerns, and conduct failure analysis. Perform risk assessments and implement mitigation strategies during product design. What We Need to See: B.Sc. in Mechanical Engineering or higher. 10&#43; overall years of relevant experience including 4&#43; years of experience managing teams of engineers in R&D environment. Proven expertise in developing, testing, and manufacturing of complex automated connection systems with precise moving parts, pneumatic and electro-mechanical systems design. Strong hands-on experience with 3D CAD tools (Creo preferred) static and dynamic mechanical simulation Solid background in designing components and sub systems an

DR ASHWANI KUMAR VIJ : Have been retained to select one : Regional Sales Manager( IT Infrastructure Services ). for a professionally managed IT Infrastructure solutions and Data Center and cloud Services organisation. Candidate Profile etc : - Should possess around 8 - 10 + Years of sales and business development experience in IT Infrastructure services and system integration with good understanding of data center and cloud services business - should be able to identify and collaborate with solution providers such as ERP, CRM, Massaging and offer complete solutions to the identified and specfic industries / customers. - should possess strong ability to identify business opportunities and develop solutions for customer requirements - strong communication and presentation skills - Post Graduates / MBA / Engineers with good understanding of IT is essential. WFM option is available. Excellent Salary Package would be offered to the right candidate. Those who have applied earlier need not apply again. DR ASHWANI KUMAR VIJ : President : Sales Recruitment Email Complete / Structured CV, with Scanned copy of Visiting Card, Salary Drawn / Expected details at :

recruitmentCRM
View job →

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of frontier AI systems. Through a combination of strategic partnerships and self-built data center campuses, we are scaling the physical infrastructure needed to deliver compute at unprecedented scale. The Commissioning organization is responsible for ensuring this infrastructure is safely tested, validated, integrated, and transitioned into reliable operations. As the portfolio grows, the team is building common standards, processes, tools, and reporting systems that allow commissioning programs to operate consistently across projects while giving teams and leadership clear visibility into readiness, risk, and execution. About the Role We are seeking a Commissioning Program Manager to build and scale the operating systems behind OpenAI’s infrastructure commissioning programs. You will own the development and continuous improvement of commissioning standards, processes, tools, dashboards, and KPIs across the infrastructure portfolio. You will work closely with commissioning and construction teams to translate field execution needs into practical playbooks, workflows, templates, metrics, and reporting mechanisms that teams can use from construction readiness through testing and turnover. This role sits at the intersection of infrastructure delivery, program management, process design, and data. The ideal candidate understands how complex construction projects operate and can turn fragmented workflows and project data into repeatable systems that improve execution without creating unnecessary administrative burden. Key Responsibilities Develop and maintain commissioning program standards, playbooks, process maps, templates, checklists, stage gates, and acceptance criteria across infrastructure projects. Establish consistent workflows for commissioning planning, construction readiness, QA/QC, issue management, document control, testing evidence

awsrestai
View job →
O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →
O
OpenAI
📍 Singapore• Full-time
1mo ago

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. We work closely with hardware design teams, manufacturing partners, suppliers, and data center operations to deliver the compute platforms that power frontier AI. The Manufacturing Engineering team ensures our hardware can be built, tested, deployed, and supported at hyperscale. We bridge Hardware Engineering, Quality, Supply Chain, Contract Manufacturers, and Hardware Operations to continuously improve manufacturing performance throughout the product lifecycle. As our fleet grows globally, sustaining manufacturing engineering becomes increasingly important to maintain product quality, improve manufacturability, and rapidly resolve production issues. About the Role We are seeking a PCBA Manufacturing Engineer (Sustaining) to support production and continuous improvement of printed circuit board assemblies (PCBAs) used throughout Industrial Compute hardware platforms. This role focuses on sustaining engineering after product launch. You'll partner closely with Hardware Design, Quality, Manufacturing, Test Engineering, Supply Chain, and our contract manufacturers to resolve production issues, improve manufacturing yield, reduce failures, implement engineering changes, and ensure stable high-volume manufacturing. The ideal candidate has experience supporting complex server, networking, accelerator, storage, or high-performance electronics manufacturing environments. Key Responsibilities Own sustaining manufacturing engineering for PCBA production across multiple hardware platforms. Drive root cause investigations for manufacturing defects, field failures, and production escapes. Partner with Hardware Design Engineers to improve manufacturability (DFM/DFA/DFT). Support engineering change orders (ECOs) and manufacturing change implementation. Work directly with contract manufacturers to improve production yield, cycle time, and quality. Analyze manuf

awsrestai
View job →

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Security Engineer to join our First-Party Hardware team. In this role, you will own the end-to-end security foundation for OpenAI's first-party AI hardware systems, working across hardware security, embedded security, system security, and practical deployment at data center scale. You will partner with silicon, hardware, firmware, infrastructure, manufacturing, operations, and security teams to define and deliver system-level device trust. This includes boot integrity, device identity, provisioning, attestation, management-plane security, storage encryption, debug controls, firmware update and recovery, RMA, and decommissioning. You will be accountable for turning threat models into requirements, requirements into implementation, and implementation into validation evidence that can support launch decisions. Location: San Francisco, CA (Hybrid: 3 days/week onsite) Relocation assistance available. In this role, you will: Own security requirements, threat models, validation strategy, and launch-readiness evidence for first-party hardware platforms from early design through production deployment. Design and review secure boot, measured boot, roots of trust, platform firmware resilience, firmware signing, recovery, and anti-rollback strategies across heterogeneous devices. Own device identity, provisioning, enrollment, attestation, certificate lifecycle, and key-management requirements across manufacturing and data center bring-up. Harden management

awsrestai
View job →

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

pythonsqlaws
View job →
O
1mo ago

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

pythonawslinux
View job →
O
1mo ago

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of AI at unprecedented scale. Through a combination of strategic partnerships and self-built campuses, we are developing and operating large-scale data center infrastructure across power, cooling, networking, compute, construction, and site operations. The scale and complexity of this infrastructure introduces a broad range of environmental, health, and safety considerations across site development, design, construction, equipment deployment, commissioning, and ongoing operations. EHS is a critical part of how we build infrastructure that is safe, resilient, compliant, and capable of operating at scale. About the Role We are seeking an EHS Lead to establish and drive environmental, health, and safety strategy across OpenAI’s rapidly expanding compute infrastructure portfolio. This role will develop the EHS framework for large-scale data center development and operations, partnering closely with engineering, construction, infrastructure delivery, facilities, operations, security, legal, environmental, and external development partners. The EHS Lead will help ensure that safety and environmental considerations are embedded into projects from early design and site development through construction, commissioning, and operations. The role will establish standards and operating mechanisms, assess and mitigate risks, oversee EHS performance across internal teams and third-party partners, and provide technical leadership on complex or high-consequence safety issues. Success requires the ability to operate strategically while maintaining strong technical depth and executional rigor in fast-moving, highly complex infrastructure environments. In this role, you will: Develop and own EHS strategy, standards, programs, and operating mechanisms across OpenAI’s data center and compute infrastructure portfolio. Establish scalable EHS requirements for site de

awsrestai
View job →
🔔

Get new data center ssd performance validation engineer jobs by email

Daily job updates · Unsubscribe anytime