Jobiba hiring network

Server Cpu Hardware Systems Lead Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current server cpu hardware systems lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

G
3 days ago

We are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute rack validation team. As a diligent leader in Systems Engineering, you will drive multiple aspects of validation throughout the life cycle of the program. In this high visibility position, you will be part of a leading team to innovate and improve system bring-up and enablement abilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The technical leader will be driving keys areas of system validation including leading first silicon & system bring-up (nodes and rack level systems) - rack level systems and blades will be based of ARM server architecture. Candidate will be immersed in challenging system enablement work, system validation (end-to-end) methodology, tests development and execution as well as triage/debug of critical issues to meet critical program milestones at POR quality. The candidate will also be a key contributor to state-of-the-art HW and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: Lead the systemenablement (including first silicon and other FW components) to ensure system capabilities are brought up as per plan of record and system architecture spec. Drive organization wide methodology for Firmware integration and best known configuration (HW/FW/SW) usage model by leading the release of deployment ready solutions. Develop key methodologies, lab HW and system SW capabilit

linuxaigo
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st

awsazurerest
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team The Stargate team is responsible for building the physical infrastructure that powers large-scale AI systems. We design and deliver next-generation data centers optimized for dense compute clusters, advanced networking, and rapidly evolving hardware platforms. This work sits at the intersection of hardware engineering, systems architecture, and infrastructure execution—translating cutting-edge compute roadmaps into scalable, production-ready environments. Our teams partner across silicon vendors, server and storage OEMs, networking teams, and data center engineering organizations to bring new capacity online quickly, reliably, and at global scale. About the Role We are seeking a CPU & Storage Technical Lead to define and drive the server compute and storage architecture strategy for Stargate infrastructure. In this role, you will own technical direction across CPU platforms, memory configurations, local and disaggregated storage systems, and their integration into large-scale AI clusters. You will evaluate vendor roadmaps, lead platform tradeoff decisions, and ensure compute and storage systems are optimized for training, inference, and supporting services. You will work cross-functionally with hardware engineering, performance modeling, networking, supply chain, and deployment teams, as well as external partners such as AMD, Intel, OEMs, ODMs, and storage vendors. This is a highly strategic role for someone who can operate deeply at the component level while also driving long-range infrastructure decisions. Key Responsibilities Own CPU and storage technical strategy for Stargate compute infrastructure across current and future generations. Evaluate CPU platforms across performance, efficiency, memory bandwidth, PCIe topology, cost, and roadmap alignment. Define storage architectures for AI environments, including boot media, local NVMe, shared storage, caching tiers, metadata services, and high-performance data pipelines. Drive server platform de

awsrestai
View job →
T
Tenstorrent
📍 AustinFull-time
3 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our Tensix Team is building the next generation of high-performance AI compute systems. We’re looking for a Power Architect to drive architectural strategy, modeling, and design decisions that shape how power is understood and optimized across our products. This is a hands-on role with massive influence over how we build power-aware systems from the ground up. This role is hybrid, based out of Santa Clara, CA, Boston, MA, Austin, TX or Toronto. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are 15+ years experience with power estimation tools like PowerArtist, PtPX, RTL Architect, and PrimePower. Skilled in modeling and optimizing power at the architectural level, with deep knowledge of power-gating, voltage domains, and leakage control. Track record of influencing architectural and micro-architectural changes that meaningfully reduced design power. Proficient in Verilog HDL, Design Compiler, C/C++ and Python. Background in power-optimization of compute datapath and/or interconnects. What We Need Predict power consumption early in architecture and track it through RTL evolution. Propose architectural changes for power optimization across server and non-

pythonawsai
View job →
T
3 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are seeking a hands-on Data Center Technician, IT Contractor to support the installation, deployment, relocation, cabling, inventory, maintenance, and decommissioning of servers, network equipment, storage systems, and other IT infrastructure. The successful candidate will follow established procedures, maintain accurate documentation, and coordinate effectively with internal technical teams and data center personnel. This role will be a 6-month contract fully on-site, based out of the downtown Toronto, ON Beanfield data center, with travel to Tenstorrent office locations and other data center facilities as required. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You bring approximately 2 to 5 years of hands-on experience in data center operations, IT infrastructure, server or network hardware, or a similar technical role. You are comfortable working with servers, network equipment, storage systems, racks, rack layouts, and copper and fiber optic cabling. You are detail-oriented, organized, and able to follow documented procedures, cabling standards, installation instructions, and safety requirements. You communicate clearly, work well wi

awsaisem
View job →
T
3 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. This role sits at the center of cutting-edge AI hardware development, keeping the servers, PCIe systems, and engineering infrastructure running that power next-generation compute. You’ll be hands-on with rapidly evolving prototype and production systems, installing, maintaining, and troubleshooting hardware in fast-paced R&D and data center environments. Acting as a critical bridge between hardware engineers, software teams, and IT, you’ll help ensure seamless access to the platforms that turn ideas into working silicon and systems. This role is onsite, based out of Toronto, Canada or Austin, Texas. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A hands-on hardware professional who enjoys building, maintaining, and troubleshooting complex computing systems. Comfortable working in fast-paced R&D environments where hardware, firmware, and software are constantly evolving. Knowledgeable in computer architecture, operating systems, and hardware diagnostics, with strong problem-solving skills. Collaborative, detail-oriented, and motivated to improve processes through documentation, scripting, and automation. What We Need Inst

awsaisem
View job →
T
3 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a Field Application Engineer to serve as the technical bridge between Tenstorrent and customers across Southeast Asia. Based in Singapore, you will work closely with customers, partners, sales, and global engineering teams to understand AI and machine learning workloads, guide technical evaluations and deployments, troubleshoot issues across hardware and software, and help customers realize the performance of Tenstorrent’s AI platforms. This is a highly visible, customer-facing role that combines hands-on technical problem solving, solution development, and regional relationship building, with regular travel throughout Southeast Asia. This role is remote, based out of Singapore. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A customer-focused technical professional who can build trust with application developers, engineering teams, and business stakeholders. Comfortable translating complex AI/ML hardware and software concepts into clear recommendations for both technical and non-technical audiences. A proactive and self-directed problem solver who can coordinate internal teams, external service providers, and customer sta

pythonawsmachine learning
View job →
M
Modal
📍 New YorkFull-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track

aigorust
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake’s cloud spend is in billions of dollars per year. Hence, it is critical for us to govern and optimize our cloud spend, both for margins and long-term competitive advantage. Cloud Efficiency team’s charter is to build scalable products that enable governance, monitoring and optimization of cloud spend. Think of this as Observability for cloud costs and efficiency. The team’s vision is to “Transform cloud spend into a competitive advantage by empowering teams to continuously optimize the per-unit cost.” In order to improve the overall cloud efficiency (i.e. cost per unit), it is critical to build monitoring products that collate costs with other factors such as utilization, attribution, hardware performance and architecture. Hence, there is an opportunity to build a unified, self-serve cloud efficiency product across Snowflake, that delivers actionable, real-time efficiency datasets through streamlined user experiences. This will enable thousands of engineers at Snowflake and will elevate cloud efficiency at Snowflake for long-term success. When developing these solutions, we think about the problem end-to-end: how do we collect data from different stacks (e.g. costs from AWS, GCP, Azure and CPU, Memory, Utilization) across Snowflake reliably, how do we store it eff

pythonjavaaws
View job →
S
Supabase
📍 RemoteFull-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. Edge Functions are server-side TypeScript functions, distributed globally at the edge - close to your users. They power use cases such as webhook receivers, AI inferences, OG image generation and real-time bots (Slack, Discord, etc). Built on top of Supabase Edge Runtime : an open-source, Deno-based runtime written in Rust that runs JavaScript, TypeScript, and WASM services. We want developers to be able to build truly global applications by distributing both compute and data globally. Infrastructure concerns like regions, cold starts, CPU and memory provisioning should fade into the background so developers can focus on iterating on business logic. We are looking for experienced and passionate engineers to help us go further in this vision. What You’ll Own Evolving Supabase Edge Runtime - an Open-sourced Rust-based host that runs the Deno isolate, manages the main/user runtime split, and enforces per-request memory and CPU limits. Implementing monitoring, alerting, and OpenTelemetry tracing across the runtime, then using that visibility to drive optimizations that improve latency and reliability of the service. Working closely with the Deno and other open-source teams, contributing to upstream and relaying our users' requirements. Participating in an on-call rotation to keep Edge Functions healthy in production. Help manage and improve features like scheduled functions, background tasks, WebSockets streaming, ephemeral file storage, and custom routing. Integrating functions more tightly with the rest of the Supabase stack - Auth, Postgres, Storage, and Realtime. Expanding functions to support more use cases (AI inference, MCP servers, hosting simple websites, URL shorteners). Improving the DX

javascripttypescriptjava
View job →
R
3 days ago

Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. AI Algorithms Engineer (PHD Holder Only) 5 Work Days Per Week Office Near Tai Seng MRT, Singapore Medical Benefits Provided Entitled to Yearly Bonus & Performance Bonus Job Requirements: PHD Holder in Computer Science, Applied Mathematics, Electrical Engineering, Pattern Recognition, Artificial Intelligence, Automatic Control, Operations Research, Biology, Physics / Quantum Computing, Neuroscience, Statistics or a related field. Familiar with common machine learning and deep learning algorithms and keeping track with the latest SOTA implementations. Strong programming skill in Python, C / C++, proficient in mathematical / statistical concepts and exceptional coding skills Hands-on experience with AI / ML frameworks be familiar such as Caffe, PyTorch, TensorFlow, MxNet etc. Have rich project experience in machine learning and deep learning, be familiar with common algorithm models, such as CNN, RNN, LSTM, Transformer, ViT, etc., and be able to improve and innovate models according to actual problems. Experience in familiar the design, parameter tuning and optimization methods of neural network models is a plus Experience in model compression and in the transplantation and optimization of deep learning forward inference on various platforms, including NPU / GPU / DSP / ARM on mobile platforms and CPU / GPU on server platforms is also a plus. Strong logical thinking and problem-solving ability, able to independen

pythonmachine learningai
View job →
R
Roblox
📍 San MateoFull-timeFrom $293.8K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Engine Networking Team pulls the players together by ensuring the communication of the game state to all. As a Principal Engineer on this team you will help the players experience the game as a nearly synchronous world. The networking and asset loading team plays a key role in a smooth experience for the players. You will work in all areas of the game platform in your quest for real-time communication of every part of Roblox. You Will: Lead engineers with 8+ years of industry experience Be experienced with one of these area: asset loading, rendering, and networking coming from a Game Engine/Studio. Be an amazing systems-level C++ programmer and be fascinated by the actual work the CPU does when you use smart pointers, templates, virtual functions, and blocks of memory, both structured and raw Have a keen to each millisecond of the network exchanges: You know where the time goes and how to reduce the waste Understand what happens on the operating system level when certain code is completed You Have: Worked on the guts of a multi-player game engine, solving problems related to scale, performance, latency, and throughput in client/server environments. Worked on a very large multithreaded d

awsgitai
View job →
N
29 days ago

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li

pythondockerkubernetes
View job →
T
Tenstorrent
📍 AustinFull-time$100K – $500K/yr
3 days ago

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a highly skilled and experienced Engineer to lead post-silicon power characterization and correlation activities for cutting-edge semiconductor products. In this role, you will be responsible for developing and executing detailed power measurement strategies on silicon, correlating results with pre-silicon models, and driving improvements across power architecture, design, and modeling methodologies. You will serve as a key technical leader, interfacing across design, architecture, validation, and systems teams to ensure silicon meets power and performance specifications under all operating conditions. This role is hybrid, based out of Toronto, ON or Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who you are A Principal-level engineer with 8+ years in silicon power analysis and characterization, and a Master’s or PhD in EE, CE, or related field. Deep understanding of digital and mixed-signal power domains, including DVFS, leakage vs. dynamic power, and power gating. Highly proficient in lab-based power measurement using oscilloscopes, current probes, power analyzers, and SMUs, plus Python/Perl/MATLAB

pythonawsgit
View job →
O
17 days ago

Technical Program Manager – Applied Infrastructure About the Team The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety. The Applied Infrastructure TPM team partners across engineering to lead foundational programs that ensure OpenAI’s infrastructure can meet current and future demand. About the Role We’re looking for a seasoned Technical Program Manager to drive critical infrastructure programs across the Applied organization. This TPM will focus on cross-cutting initiatives such as general compute capacity planning, process transformation, cost and quota attribution and optimization, and coordination across infrastructure and product stakeholders. There will also be focus on evolving OpenAI’s infrastructure to support growth, scale and new products. This work is core to how OpenAI manages and grows its infrastructure footprint in a disciplined, scalable way. Location: San Francisco, CA (Hybrid – 3 days/week in-office) In this role, you will: Serve as the DRI for complex infrastructure programs spanning CPU planning, orchestration, and other resource management domains (e.g. networking, storage). Build and operationalize systems to capture demand signals, model future capacity needs, and align infrastructure planning across internal teams and partners external to the company. Partner closely with Infrastructure, Product and Finance teams to forecast infrastructure usage patterns and ensure supply/demand alignment. Lead cost attribution and quota enforcement programs to promote stability and ensure equitable access to resources across teams. Drive simplification and standardization of infrastructure tooling and processes across Applied and Infra organizations. Drive cross functional programs to evolve our infrastructure to support new growth and scale Work with external v

awsazurerest
View job →
🔔

Get new server cpu hardware systems lead jobs by email

Daily job updates · Unsubscribe anytime