About Us Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Job Summary Working within the logical design team, the silicon logical design engineer is responsible for a wide range of logical design tasks. The team is responsible for delivering the Microarchitecture and RTL design to implement the chip architecture specification for Graphcore Silicon, working closely with other engineers within the Silicon team. The successful candidate will be responsible for helping the team deliver high quality micro-architecture and RTL for Graphcore chips, working within the logical design team and with the broader Silicon team to ensure we meet the company objectives for Silicon delivery. Responsibilities and Duties Integrate IP and subsystems into top-level SoC designs Develop and maintain build and configuration environments Perform synthesis, linting, CDC/RDC, and timing checks at the SoC level Support verification and physical design teams through clean interface hand-offs Debug and resolve integration-related issues across multiple hierarchies Contribute to the continuous improvement of integration flows and automation Producing high quality microarchitecture and other documentation Ensure good communication between sites to maintain consistent working practises Candidate Profile Essential skills: Logical design experience in relevant industry Experience range 4-8 years in Semiconductor Industry/Product development exposure. Be highly motivated, a self-starter, and a team player Ability to work across teams and debugging issues seen to find root caus
Jobiba hiring network
Data Center Ssd Performance Validation Engineer Jobs
8,157 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current data center ssd performance validation engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
NVIDIA is well positioned as the 'AI Computing Company', our GPUs being the brains that power modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We are looking for a Senior Software QA Test Development Engineer to join in the mission of crafting a distributed technology for all NVIDIA teams that remotely manage 10s of 1000s of resources in a simple and controlled fashion, allowing engineers to focus on engineering and automation, rather than being burdened by manual operational tasks. SWQA test developer engineers at NVIDIA are responsible for creating test plans, execution, and reporting, as well as developing scripts for test automation, designing and developing tools for the QA team, and developing integration tests for validation. As a test developer, you must identify weak spots and constantly design better and more creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. The ideal candidate must have strong programming skills and hands-on experience using AI development tools to improve quality and productivity across the end-to-end QA workflow. This includes leveraging AI assistants for test automation, code generation, debugging, and enhancing testing efficiency. During the interview process, we will assess your ability to effectively use AI development tools and evaluate your programming capabilities to ensure you can deliver high-quality solutions. What you’ll be doing: Architect, implement, and evolve scalable agentic end to end SWQA workflow, automated test frameworks, infrastructure, and tooling for complex software products. Define test strategy and quality gates across functional, integration, regression, reliability, and release-validation workflows. Build and maintain high-value automated coverage for Linux-based, co
NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer! What you'll be doing: Primary responsibilities will include building AI/HPC infrastructure for new and existing customers. Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting. Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement. Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields. At least 5+ years of professional experience in networking fundamentals, Ethernet or InfiniBand World. Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS, etc. Possess solid working knowl
Associate Lead - EHS will be responsible for implementation, and management of Environmental, Health, and Safety (EHS) policies, procedures, and programs across all Data Centre projects. This role involves ensuring a safe and compliant work environment through regular audits, risk assessments, and training initiatives. Source: Adani Group | Job ID: 53470
Associate Lead - EHS will be responsible for implementation, and management of Environmental, Health, and Safety (EHS) policies, procedures, and programs across all Data Centre projects. This role involves ensuring a safe and compliant work environment through regular audits, risk assessments, and training initiatives. Source: Adani Group | Job ID: 48429
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
The Lead - Billing (MEP) will be responsible for end-to-end billing, quantity surveying, contract administration, and commercial management for Mechanical, Electrical, Plumbing (MEP), Fire Fighting, HVAC, ELV, and Utility works. The role involves preparation and certification of RA bills, contractor and subcontractor billing, client billing, variation management, claims administration, BOQ reconciliation, and cost monitoring for large-scale construction projects including high-rise buildings, industrial plants, data centers, commercial developments, and infrastructure projects. Source: Adani Group | Job ID: 59172
We are looking for an enthusiastic software engineer to join our AI networking acceleration team, to work on a groundbreaking open-source library, using hardware offloads, GPU Kernels and RDMA network cards. Our product is a performance-oriented low-level infrastructure, crafted to change the way inference works. We thrive as a team in a deeply strong environment, and we're passionate about innovation. The rewards are sweet and include working with some of the brightest people in the industry, an aggressive compensation plan that rewards top performers, and the opportunity to collaborate on products that transform daily the way people work and play. What you'll be doing: Developing a highly optimized inference framework Running on the world’s largest supercomputers and data centers. The work environment is dynamic and challenging as our employees work on innovative, next-generation products at the forefront of technology in terms of performance, scalability, and features. What we need to see: B.Sc. or equivalent experience in Computer Science or Software Engineering 6+ years of experience in modern C++ / C / Rust development 3 years of experience in Linux environment and familiarity with development tools Deep knowledge of the TCP/IP network stack Understanding of computer architecture and operating systems concepts Ways to stand out from the crowd: Background in Linux internals and low-level software optimizations (benchmarking, bottleneck research, performance tuning) Experience in programming CUDA kernels is an advantage Familiarity with ML frameworks and LLMs Background in parallel programming / high-performance computing / RDMA t
Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As a NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA is looking for an AI Computer Engineer to join its NVIDIA Infrastructure Specialists (NVIS) team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI factory systems in the world! NVIDIA is looking for someone with the ability to work on a dynamic customer-focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyse, define and implement large scale AI Factory projects. The scope of these efforts includes a combination of Networking, System Design and Automation while being the face to the customer. What you will be doing: Primary responsibilities will include deploying, managing, and maintaining AI infrastructure for new and existing customers. Be the domain expert with customers during planning calls through implementation. Handover-related documentation and perform knowledge transfers required to support customers as they begin rolling out some of the most sophisticated systems in the world! Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements. What we need to see: Bachelor’s degree in computer science, Electrical Engineering, or a related field, or equivalent experience.
Are you a person who likes to work in a fast-paced organization? NVIDIA is the world leader in Visual Computing. We are passionate about four markets: Gaming, Automotive, Enterprise Graphics and HPC/Cloud Datacenters; in addition to our traditional OEM business. We are well positioned as the ‘AI Computing Company’, and our GPUs are the brains powering modern Deep Learning software frameworks, accelerated analytics, big data, modern data centers, smart cities, and driving autonomous vehicles. We have some of the most forward-thinking and talented people on the planet working for us. If you're forward-thinking, hardworking, driven and if working with extraordinary people across countries sounds interesting, this job is for you. We are now looking for a Human Resources Business Partner to provide HR support onsite in Santa Clara, CA for our Worldwide Field Organization in a dynamic and collaborative environment. This is a global organization, and we are looking for someone to be passionate about supporting and building strategies to enable NVIDIA to achieve success. You’ll partner with a cross-functional group of subject matter experts to design and execute strategies for how we staff, onboard, develop, motivate, retain and organize work. You will need excellent communication skills, critical thinking and planning ability, and the agility to function in a fast paced and innovative environment. What you'll be doing: This position will be an integral enabler of the mission of our Field organization. In this position you will work with the senior leaders and leadership teams within NVIDIA organizations to develop and execute the HR strategies that champion organizational and people effectiveness. You will think strategically as well as roll up your sleeves and dive deep into practical application. You must understand business priorities and translate them into an HR ag
About Backblaze Backblaze provides reliable, high-availability cloud storage trusted by consumers, SMBs, enterprises, and developers in over 150 countries. Backblaze B2 Cloud Storage supports data-intensive workloads including backup, media, analytics, and modern AI pipelines. Our teams focus on building secure, durable, scalable systems with a strong emphasis on application security from day one and protecting customers to keep them safe. While there is a lot to celebrate in our past, there are greater opportunities ahead of us. About the Role We are seeking a Back-End Engineer for our Computer Backup Team. This team focuses on the services supporting our Computer Backup product. In this role, you will build and maintain systems to back up and restore exabytes of data. We are looking for an engineer who is excited to learn to develop high-performance, highly multithreaded architectures. We value versatility and are looking for someone comfortable designing and working across any technology as our needs evolve. Responsibilities Design and develop highly scalable, performant services and APIs in a Java ecosystem. Integrate AI coding agents and frameworks into your workflow to accelerate development, refactor legacy components, and improve code quality. Collaborate on the full technical lifecycle, including estimation, design, development, and delivery. Collaborate across the organization to ensure project success. You will often engage with QA, Data Centers, Marketing, Product Management, Support, Legal/Compliance, and Security to gather necessary context and requirements to implement solutions correctly, help customers, and solve problems. Diagnose complex issues and maintain the stability of large-scale production systems. Collaborate across functional teams and provide technical leadership to foster engineering excellence. Required Qualifications A broad track record of project deliveries and successfully shipping software systems. 3+ years of experience with Java
Location: San Francisco, CA (Remote/Hybrid Available) What is Verse? The race to AI has become the race to power. Every breakthrough in artificial intelligence depends on one thing: access to electricity. But across the country, aging grid infrastructure and years-long interconnection queues are slowing the deployment of the data centers that will power the next generation of innovation. Solving this challenge isn't just about energy—it's about unlocking the future of AI. At Verse, we're building the energy intelligence platform for the AI economy. Our software helps the world's largest energy consumers achieve faster, cheaper, and cleaner power by combining real-time control of energy assets with complete visibility into their energy portfolio. Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale batteries, energy markets, and enterprise software, we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud-native applications that scale to trillions of data points collected from electricity markets globally. You will be a part of a dynamic, robust team primarily supporting the backend needs of our Aria software product spanning hundreds of data sources, sinks, services, and jobs. Your expertise will not only have a direct impact on product decisions, but you also be well-positioned to drive the development and trajectory of our entire platform and infrastructure and influence important architectural decisions that affect the whole organization. Key Responsibilities Foster a culture and mindset of well-designed systems, test-driven software, and transparent communication with a high caliber of mutual respect and consideration for stakeholders Read and write a lot of Go, Python, and Protobuf Build, test, debug, maint
Get new data center ssd performance validation engineer jobs by email
Daily job updates · Unsubscribe anytime