←Jobiba.Me
I
Active1mo ago

HPC Dev Ops Engineer

Intel·📍 3 Locations

Employment

FULL TIME

Work mode

On-site

Experience

All levels

Salary

Not disclosed

Salary not disclosed

Check market pay for comparable HPC Dev Ops Engineer roles before applying.

Salary →

Role overview

Job description

HPC Dev Ops Engineer — 3 Locations. Apply via Workday.

I

Hiring company

Intel

Explore this employer's active roles, salary signals and company profile on Jobiba.

Keep exploring

Similar active roles

Fresh roles matched to this title and market.

View all →
G
📍 Austin, Texas, United States· Full-time

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

PythonAIGoDevOps
B
📍 San Francisco, California, United States· Full-time

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE OPPORTUNITY We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work. RESPONSIBILITIES Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving). DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation. Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis. System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack. Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance. Continuous Integration : Auto

PythonCI/CDGitRest
E
📍 Dubai, United Arab Emirates· Full-time

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Powering the METCA AI & Sovereign Cloud Revolution As Everpure , we are transcending traditional storage to deliver the Enterprise Data Cloud - a unified architecture engineered to fuel the world’s most ambitious AI, Deep Learning, and HPC projects. The METCA region is a global epicenter for AI infrastructure investment, characterised by massive capital flowing into Tier-1 sovereign GPU clouds (CSPs, Neo-scalers) and enterprise AI factories. We are seeking a Battle-Trained, High-Conviction Hunting Systems Engineer (SE) to serve as our technical tip of the spear. This is not a passive, box-pushing relationship management role. You will partner aggressively with an Enterprise Account Executive to target, break into, and land the largest AI infrastructure projects in the market, displacing legacy architectures and securing net-new footprints. WHAT YOU'LL DO Execute High-Impact Hunting: Partner closely with Account Executives to actively map out and break into net-new enterprise accounts, sovereign GPU clouds, and high-performance computing clusters. Architect the AI Factory: Design high-performance, multi-tenant data pipelines. Move beyond basic storage architecture to design full-stack environments, optimising how data nodes interact within massive GPU fabrics. Drive Technical Consensus: Lead deep-dive architectural workshops with customer GPU cluster architects while simultaneously translating complex en

AWSKubernetesRestAI
N
📍 Remote, Poland· Remote

Join our multidisciplinary team and help build and improve GPU and CPU accelerated data processing software libraries. Projects like DALI or nvImageCodec are used in all kinds of processing workflows and support NVIDIA's vision and growth. Starting from powering AI, data analytics, image processing, computer vision, and scientific simulations for leading commercial and academic organizations worldwide. In this role, you will design, develop, and optimize pioneering algorithms. Ideal candidates will have experience with accelerated computing and a passion for advancing the state-of-the-art in various computing domains. If this sounds exciting, we would love to meet you! What you’ll be doing: Developing scalable library software using modern tools and languages for various numerical method. Performance tuning, optimization, and benchmarking of algorithms on various architectures. Working closely with leadership team and other internal and external partners to understand feature and performance requirements and contribute to the technical roadmaps of libraries. Providing technical leadership and guidance to library engineers working with you. Find opportunities to improve user experience and library performance. What we need to see: PhD or MSc’s degree in Computational Science, Computer Science, Applied Math, or related science or engineering field of study is preferred (or equivalent experience). 5+ years experience developing, debugging, and optimizing high-performance parallel numerical applications on modern computing platforms, with GPU acceleration using CUDA. C/C++ programming and software development skills. Proven experience in leading and completing software development projects. Strong collaboration, communicati

PythonAIProject Management

NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,

LinuxArtificial IntelligenceAI

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The internal infrastructure team is responsible for building world-class infrastructure and tools used to train, evaluate and serve Cohere's foundational models. By joining our team, you will work in close collaboration with AI researchers to support their AI workload needs on the cutting edge, with a strong focus on stability, scalability, and observability. You will be responsible for building and operating superclusters across multiple clouds. Your work will directly accelerate the development of industry-leading AI models that power Cohere's platform North. Please Note: All of our infrastructure roles require participating in a 24x7 on-call rotation, where you are compensated for your on-call schedule. As a Staff Software Engineer, you will: Build and scale ML-optimized HPC infrastructure : Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads. Optimize for AI/ML training : Collaborate with cloud providers to fine-tune infrastructure for cost efficiency, reliability, and performance , leveraging technologies like R

PythonKubernetesGitLinux

🔔 Get job alerts

New HPC Dev Ops Engineer jobs in 3 Locations, straight to your inbox.

No spam · Unsubscribe anytime