HPC Dev Ops Engineer — 3 Locations. Apply via Workday.
Jobiba hiring network
Hpc Dev Ops Engineer Jobs
43 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current hpc dev ops engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto
Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Powering the METCA AI & Sovereign Cloud Revolution As Everpure , we are transcending traditional storage to deliver the Enterprise Data Cloud - a unified architecture engineered to fuel the world’s most ambitious AI, Deep Learning, and HPC projects. The METCA region is a global epicenter for AI infrastructure investment, characterised by massive capital flowing into Tier-1 sovereign GPU clouds (CSPs, Neo-scalers) and enterprise AI factories. We are seeking a Battle-Trained, High-Conviction Hunting Systems Engineer (SE) to serve as our technical tip of the spear. This is not a passive, box-pushing relationship management role. You will partner aggressively with an Enterprise Account Executive to target, break into, and land the largest AI infrastructure projects in the market, displacing legacy architectures and securing net-new footprints. WHAT YOU'LL DO Execute High-Impact Hunting: Partner closely with Account Executives to actively map out and break into net-new enterprise accounts, sovereign GPU clouds, and high-performance computing clusters. Architect the AI Factory: Design high-performance, multi-tenant data pipelines. Move beyond basic storage architecture to design full-stack environments, optimising how data nodes interact within massive GPU fabrics. Drive Technical Consensus: Lead deep-dive architectural workshops with customer GPU cluster architects while simultaneously translating complex en
Senior HPC Architect, Automation and At-Scale Deployment — 13 Locations. Apply via Workday.
Join our multidisciplinary team and help build and improve GPU and CPU accelerated data processing software libraries. Projects like DALI or nvImageCodec are used in all kinds of processing workflows and support NVIDIA's vision and growth. Starting from powering AI, data analytics, image processing, computer vision, and scientific simulations for leading commercial and academic organizations worldwide. In this role, you will design, develop, and optimize pioneering algorithms. Ideal candidates will have experience with accelerated computing and a passion for advancing the state-of-the-art in various computing domains. If this sounds exciting, we would love to meet you! What you’ll be doing: Developing scalable library software using modern tools and languages for various numerical method. Performance tuning, optimization, and benchmarking of algorithms on various architectures. Working closely with leadership team and other internal and external partners to understand feature and performance requirements and contribute to the technical roadmaps of libraries. Providing technical leadership and guidance to library engineers working with you. Find opportunities to improve user experience and library performance. What we need to see: PhD or MSc’s degree in Computational Science, Computer Science, Applied Math, or related science or engineering field of study is preferred (or equivalent experience). 5+ years experience developing, debugging, and optimizing high-performance parallel numerical applications on modern computing platforms, with GPU acceleration using CUDA. C/C++ programming and software development skills. Proven experience in leading and completing software development projects. Strong collaboration, communicati
NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. Our SA team is more focusing to bring NVIDIA new technology into difference industries. We help to design the architecture of AI computing platform, analysis the AI and HPC applications to deliver our value to customers, focusing on defining and solving computational challenges in LLM inference and training acceleration, as well as network communication and data transfer optimization. What You'll Be Doing: Contribute to the development of open-source inference frameworks such as SGLang and vLLM, including feature and operator development, performance optimization, and model support, in collaboration with the community. Develop and optimize KV cache offloading frameworks for LLM workloads, supporting multi-level cache offloading and reuse across CPU, SSD, and remote storage to improve inference efficiency. (Team project: FlexKV) Drive R&D on compute performance in distributed training, and explore methods and technologies for performance optimization. Study computational challenges in machine learning systems, identify common needs and bottlenecks, and build example code, acceleration libraries, or frameworks accordingly. What We Need to See: Over 5 years working experience in the technology industry, with master’s degree or above in computer science, mathematics, electrical engineering, automation, or related fields. Strong interest in accelerated computing, parallel computing, and heterogeneous computing, with the motivation to explore these areas in depth. Solid programming skills, with a good understanding of data structures and computer systems fundamentals. Strong learning agil
Are you a person who likes to work in a fast-paced organization? NVIDIA is the world leader in Visual Computing. We are passionate about four markets: Gaming, Automotive, Enterprise Graphics and HPC/Cloud Datacenters; in addition to our traditional OEM business. We are well positioned as the ‘AI Computing Company’, and our GPUs are the brains powering modern Deep Learning software frameworks, accelerated analytics, big data, modern data centers, smart cities, and driving autonomous vehicles. We have some of the most forward-thinking and talented people on the planet working for us. If you're forward-thinking, hardworking, driven and if working with extraordinary people across countries sounds interesting, this job is for you. We are now looking for a Human Resources Business Partner to provide HR support onsite in Santa Clara, CA for our Worldwide Field Organization in a dynamic and collaborative environment. This is a global organization, and we are looking for someone to be passionate about supporting and building strategies to enable NVIDIA to achieve success. You’ll partner with a cross-functional group of subject matter experts to design and execute strategies for how we staff, onboard, develop, motivate, retain and organize work. You will need excellent communication skills, critical thinking and planning ability, and the agility to function in a fast paced and innovative environment. What you'll be doing: This position will be an integral enabler of the mission of our Field organization. In this position you will work with the senior leaders and leadership teams within NVIDIA organizations to develop and execute the HR strategies that champion organizational and people effectiveness. You will think strategically as well as roll up your sleeves and dive deep into practical application. You must understand business priorities and translate them into an HR ag
We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver communication libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. DL and HPC applications have a huge compute demand already and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! We are looking for a technical leader to manage our NVSHMEM and UCX libraries. This is an outstanding opportunity to push the limits on the state-of-the-art and deliver platforms the world has never seen before. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Lead, mentor, and grow your library engineering team and be responsible for the planning and execution of projects as well as the quality, and performance of your libraries. This is a technical leadership role so you will participate in feature design and implementation. Interact with internal and external partners and researchers to understand their use cases and requirements. Collaborate with engineering teams, program and product management, and partners to define the product roadmap. Continuously review and identify improvement opportunities in established processes, infrastructure, and practices to ensure the teams are executing in the most efficient and transparent manner. What we need to see: 10+ overall years of experience in the software industry with specialization in HPC networking or system software. 4+ years of management experience. BS, MS, or Ph.D. in C
NVIDIA is the world leader in GPU Computing. We are passionate about markets including gaming, automotive, professional vision, HPC, datacenters and networking in addition to our traditional OEM business. NVIDIA is also well positioned as the ‘AI Computing Company’, and NVIDIA GPUs are the brains powering modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We have some of the most experienced and dedicated people in the world working for us. If you are dedicated, forward-thinking, and if working with hard-working technical people across countries sounds exciting, this job is for you. We are now looking for a Software QA Test Development Engineer, you will collaborate with multi-functional groups. SWQA test developer engineer at NVIDIA is responsible for test planning, execution, and reporting, you will also write scripts to automate testing, design and develop tools for QA team, or develop integration tests for validation, so QA engineer can improve productivity or optimize test plan. As a SWQA test developer, you must identify weak spots and constantly design better and creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. What You’ll Be Doing Analyze requirements and design test matrices covering functionality, performance, and edge cases. Develop test plans and cases; build and maintain automated test suites (API / UI / CLI / E2E) in Python. Leverage AI-powered tools and agentic workflows to accelerate test generation, triage, and root cause analysis. Manage the full bug lifecycle — filing, reproduction, and driving multi-functional collaboration to resolution. Reproduce and verify customer-reported issues to ensure quality before release.
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. NVIDIA GH200 superchip provides performance and productivity required for strong scaling for HPC and generative AI workload. Scale out is inherent to design of this massive superchip. We are looking for expert engineers to come and help design rack level solutions for next generation scaling AI supercomputing platforms. We are looking for a strong technical architect to own end to end manageability architecture for these products in data centers. You will work with various component leads internally and externally, drive customer use cases, align architecture with customer requirements and release best products to market. Join us at the forefront of technological advancement. What you’ll be doing: Drive server management for large clusters and data centers deploying GPUs and Grace solution from Nvidia. Work with data center architects and cloud customers to narrow down on requirements for implementation to ensure speed of light product development. Work with internal teams to make sure requirements are designed and implemented in right way with each firmware and software module Collaborate with other leads to design & build data center health management workflow. Drive reliability and optimization in firmware architecture from a data center view point. Work closely with cluster bring up team and resolve is
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. As our TT-Distributed Software Engineer, you will develop and optimize distributed software systems that power the most efficient and highest-performing AI and HPC clusters. In this role, you'll work on distributed programming across multiple nodes, utilizing systems programming, inter-node communication, and Tenstorrent’s scalable architectures to advance the state-of-the-art distributed inference and training infrastructure. This role is hybrid, based out of Santa Clara, CA; Austin, TX; or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Strong C or C++ engineer with solid foundations in systems programming, operating systems, and distributed systems principles. Enthusiastic about distributed computing, including IPC, socket programming, and cluster resource coordination. Comfortable reasoning about scalability, fault tolerance, and performance across multi-node environments. Curious and first-principles thinker who challenges conventional approaches to distributed system design. Motivated to grow into a deep technical expert in large-scale distributed AI infrastructure. What We Need Architect, implement, and optim
Get new hpc dev ops engineer jobs by email
Daily job updates · Unsubscribe anytime