By submitting your resume, you’re expressing interest in our 2027 RDSS (Research and Development Substitute Services) program. Please confirm your eligibility with the local district office before applying the role. NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on usecase power estimation, power and performance optimization What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major Software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Experience in working with the Linux and open-source software communities Understanding of the Linux power management features (scheduler, dynamic frequency scaling, runtime power management, su
Jobiba hiring network
Hardware Operations Engineer Jobs
1,283 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current hardware operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Are you ready to do your life’s work at the heart of the autonomous revolution? NVIDIA’s SWQA organization is seeking a world-class Software QA Test and Tool Developer to join our Automotive Platform team, where the code you validate ensures the safety of millions on the road. In this role, you won't just be testing software; you will be architecting the security and reliability of the next generation of intelligent vehicles. We are looking for engineers who are as comfortable navigating low-level product architecture as they are deep-diving into complex product use cases with passion for quality. This is a high-impact, hands-on position focused on our industry-leading automotive products, offering a rare opportunity to influence the core of our tech stack. You will also build the tools and frameworks that define performance standards for systems running on Linux and QNX. What you’ll be doing: Design, execute, and automate comprehensive test cases and test scenarios to validate our automotive platforms using various test methodologies to identify and track actionable defects and track them to closure. Participate in deep-dive reviews of product requirements and technical designs, providing critical feedback to ensure features are built for testability and security from day one. Partner closely with project management, hardware teams, and software developers to provide rigorous technical analysis of bugs and publish data-driven statistical reports for global team members. Architect and maintain a distributed test automation framework capable of managing high-concurrency workloads across an extensive automation farm of hundreds of concurrent systems. Develop sophisticated test libraries and automation solutions to accelerate development cycles and expand automated test coverage for re
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning. What you will be doing: Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters. Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components. Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware. What we need to see: M.S./Ph.D. degree in CS/CE or equivalent experience. 5+ years of relevant experience. Excellent C/C++ programming and debugging skills. Strong experience with Linux. Expert understanding of computer syst
About BlockTech BlockTech is a fast-paced algorithmic trading firm facilitating global cryptocurrency derivatives and spot trading while expanding into new markets. As we continue to grow rapidly, we are looking for a Software Engineer to join our Foundation team amid our exciting scale-up phase! You will Build & optimize: Design, develop, and maintain high-reliability, low-latency, and high-throughput foundational systems that enable our trading and technology teams to scale efficiently. Ingest & aggregate: Collect trading business data with minimal latency impact and ingest both public and private exchange information into our trading system. Store & stream: Develop and maintain infrastructure for real-time data aggregation and long-term storage, as well as our Kafka-based messaging systems. Collaborate & support: Work closely with multiple teams, assisting them in integrating with and making the most of our foundational systems. Innovate: Drive projects from concept to deployment with full ownership, and explore new tools, frameworks, and approaches to keep our infrastructure best-in-class. The Foundation team develops core software infrastructure (libraries, frameworks, and systems) for BlockTech, solving common problems and lending its expertise to enable other teams to stay focused on their respective domains. They own, develop, and configure a wide variety of critical, high-reliability software, ranging from low-level ultra-low-latency shared memory IPC libraries to high-throughput data buses and data aggregation systems including Kafka, NATS, PostgreSQL, and Iceberg. They work primarily in Rust, but also use Python and SQL. If you thrive on low-level problem-solving, building robust frameworks from scratch, and enabling others to move faster, this role is for you. What We're Looking For Essential: 5+ years of experience as a Software Engineer, with a strong focus on systems-level optimisation and awareness of hardware constraints Proficiency
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role As a Corporate it engineer at WRITER, you'll be the backbone of our internal technology ecosystem. You'll design, build, and scale the systems that empower our rapidly growing team to do their best work. This isn't just about closing tickets; it's about architecting seamless, secure, and automated experiences across our entire SaaS and hardware landscape. You'll have the opportunity to shape how a leading AI company operates from the inside out, ensuring our infrastructure scales as fast as our ambitions. This role is hybrid and based out of our San Francisco or New York City office hubs. You will report to Staff IT lead 🦸🏻♀️ What you'll do Own the architecture and administration of our core enterprise SaaS platforms — including Google Workspace, Slack, and Okta Build robust automations and API integrations using Python and Terraform to eliminate manual IT processes and engineer zero-touch onboarding and offboarding for our global workforce Drive our identity and acc
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role As a Corporate it engineer at WRITER, you'll be the backbone of our internal technology ecosystem. You'll design, build, and scale the systems that empower our rapidly growing team to do their best work. This isn't just about closing tickets; it's about architecting seamless, secure, and automated experiences across our entire SaaS and hardware landscape. You'll have the opportunity to shape how a leading AI company operates from the inside out, ensuring our infrastructure scales as fast as our ambitions. This role is hybrid and based out of our London office hub. You will report to Staff IT lead 🦸🏻♀️ What you'll do Own the architecture and administration of our core enterprise SaaS platforms — including Google Workspace, Slack, and Okta Build robust automations and API integrations using Python and Terraform to eliminate manual IT processes and engineer zero-touch onboarding and offboarding for our global workforce Drive our identity and access management strategy b
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including: model architecture and MoE routing optimization decoding and inference-time algorithm improvements software/hardware co-design for GPU acceleration performance optimization without compromising model quality Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production. You may be a good fit
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs. If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact. What You’ll Work On Build and own the training framework responsible for large-scale LLM training. Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing). Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100). Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics. Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training. Investigate and res
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
What you’ll do Own and evolve the Quality Management System (QMS) to support a regulated medical device development program, including design controls and DHF maintenance. Establish and enforce requirements traceability: user needs → design requirements → verification/validation artifacts and change control. Define and run the program-level V&V strategy (verification, validation, and test coverage), including test plans, protocols, reports, and acceptance criteria. Drive risk management activities (e.g., DFMEA / PFMEA, hazard analyses) and ensure mitigations are reflected in requirements and verification. Lead document control: reviews, approvals, training, retention, and audit readiness. Partner with engineering to make quality “native” to the dev workflow (automated testing, release gates, software configuration management). Prepare the program for audits and inspections, including hands-on audit leadership. What we’re looking for Senior experience leading quality for complex hardware + software products in a regulated environment. Deep familiarity with design controls, DHF, document control, risk management, and verification planning. Strong systems thinking and the ability to translate ambiguous product intent into testable requirements. Comfortable collaborating directly with multidisciplinary engineering (recon/ML, embedded, mechanical, EE, cloud). Useful experience Regulated product quality leadership (ISO 13485 / 21 CFR 820 or equivalent), including audit readiness and FDA-facing work. eQMS + document control fluency (e.g., Greenlight Guru) that integrates cleanly with modern engineering workflows.
What you’ll do Be the generalist EE for the scanner system: integration, bring-up, debugging, and making the electrical side of the device reliable and serviceable. Own ultrasound experimentations that feeds the image reconstruction team Design and execute experiment setups for transducer characterization (element sensitivity, bandwidth, cross-talk mapping, beam profile measurements) and ex vivo / phantom clinical testing. Acquire, process, and analyze RF and baseband signals for data quality assessment and benchmarking. Design simple boards and adapters as needed (monitoring, power/safety, interface/conditioning), and take them from prototype through a stable revision. Prototype quickly, then harden what works: wiring/harnessing, grounding, safety interlocks, and reliable integration across subsystems. Own practical test setups and documentation (fixtures, scripts, procedures) that make experiments repeatable and results comparable over time. What we’re looking for Strong hands-on EE background with experience building, debugging, and iterating on real systems in the lab. Solid understanding of signal processing fundamentals — knows what to measure, how to condition and digitize it, and how to evaluate signal quality in the context of an imaging system (SNR, bandwidth, dynamic range, artifacts). Comfortable spanning system integration + occasional design work (schematics/layout reviews or light PCB design) in a fast-moving environment. Ability to work at the boundary between hardware and algorithms: measure reality, communicate constraints, and help close gaps vs simulation. High agency and practicality: able to set up experiments, get trustworthy data, and unblock others on a lean team. Useful experience Analog/mixed-signal, or high-speed data capture experience; strong instincts for instrumentation and noise/debugging. Ultrasound or acoustic sensor handling: hydrophone calibration and field mapping, transducer impedance characterization, element-level sensitivity
What you’ll do Act as the in-house electrical lead for Midjourney Medical: own the electrical architecture of the scanner and the technical direction for all board-level design. Own complex board design end-to-end: architecture, schematic capture, layout (high-speed digital, analog/mixed-signal, power), DFM/DFT, fabrication and assembly vendor management, bring-up, and revision control. Write firmware for embedded targets (MCU/SoC): drivers, real-time control loops, safety-relevant logic, bootloaders, and field update paths. Audit and update HDL (FPGA) code for high-throughput data acquisition, timing/synchronization, triggering, and pre-processing of ultrasound and sensor data streams. Define electrical interfaces and data contracts with software, recon/ML and mechanical teams: timing budgets, clocking/sync, signal integrity, connectors/harnessing, and failure modes. Establish electrical engineering rigor: design reviews, schematic/layout review checklists, bring-up procedures, test fixtures, and documentation suitable for a regulated medical device program (DHF, traceability, change control). Mentor and grow the electrical function; select and manage external design partners where leverage is high. What we’re looking for Deep experience designing complex boards from blank page to stable revision, including high-speed digital and analog/mixed-signal domains. Strong schematic and layout skills (Altium/KiCad or equivalent) with real signal integrity, power integrity, grounding, and EMI/EMC instincts. Solid embedded firmware background in C/C++ (and Python for tooling): peripherals, DMA, interrupts, real-time constraints, and debugging on hardware. Practical HDL experience (VHDL/Verilog/SystemVerilog) for data acquisition, timing, and streaming interfaces. Track record of owning bring-up and debug on real hardware: scopes, logic analyzers, and disciplined root-cause analysis. Technical leadership: clear trade-offs, strong written documentation, and the ability to set
About Flexport: At Flexport, we believe global trade can move the human race forward. That’s why it’s our mission to make global commerce so easy there will be more of it. We’re shaping the future of a $10T industry with solutions powered by innovative technology and exceptional people. Today, companies of all sizes—from emerging brands to Fortune 500s—use Flexport technology to move more than $19B of merchandise across 112 countries a year. The recent global supply chain crisis has put Flexport center stage as we continue to play a pivotal role in how goods move around the world. We are proud to have the support of the best investors in the game who believe in our mission, solutions and people. Ready to tackle global challenges that impact business, society, and the environment? Come join us. Help Win New Business The Opportunity: We are scaling our dedicated Data Center practice, and we are looking for the person who will lead it. This is a founding commercial role. You will start as a team of one, owning the full sales motion end-to-end, and you will build the team around you as the practice grows. You will define how Flexport goes to market with hyperscalers, hardware OEMs, and the broader ecosystem. You will set the playbook, win the first marquee accounts, and hire the people who scale what you build. Reporting directly to the Regional General Manager, you will operate with a high degree of autonomy and direct access to executive leadership. This role features a 50/50 compensation model (Base + Uncapped bonus, with accelerators), with OTEs of $350k+, designed for strong leaders who take bets on themselves. Why This Role Is Different: The data center logistics market is not a standard freight problem. A single AI compute rack can cost more than $1 million and weigh up to 4,000 pounds. Racks contain Class 9 dangerous goods (lithium-ion batteries) and liquid cooling systems requiring specialized handling. Construction sequencing failures delayed 57% o
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Devices Platform team's mandate is to lay the foundation of Nuro's onboard software for our sensor and compute platform, including device drivers, inter-device protocols and pipelines, and device runtime APIs. Sensors and compute hardware are the eyes, ears, and brains of our self-driving robots. We are creating the hardware-agnostic platform to be used by the perception and autonomy SW stack, and to realize the full potential of our sensor and compute HW in reliability, quality, and performance. The projects we work on are high impact and high visibility within Nuro. This team is also responsible for working with internal stakeholders and external suppliers to define, evaluate, integrate the next generation HW platform for Nuro's products and to build the necessary tooling to assist continuous testing and validation. About the Work Design and develop sensor and compute systems for robotics Architect and/or deploy Nuro
Get new hardware operations engineer jobs by email
Daily job updates · Unsubscribe anytime