Jobiba hiring network

Hardware Operations Engineer Jobs

1,283 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current hardware operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

E
22 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are hiring a senior quality engineer to own System Testing for Pure’s FlashArray products. You will validate stability, resiliency, and performance under sustained, production‑like workloads , far beyond basic functional testing. You will design and run multi‑week, customer‑like scenarios that combine workloads, failovers, upgrades, and fault injections, and use the insights to influence architecture, design, and release decisions. This is a hands‑on, high‑impact role at the intersection of architecture, systems, and large‑scale testing—acting as a key quality gate before releases reach Pure’s customers. WHAT YOU'LL DO Own System Testing strategy for releases Define System Testing strategy and test plans for major features and releases, focusing on stability, longevity, and end‑to‑end behavior, not just feature correctness. Design realistic, high‑value scenarios Build scenarios that mirror Pure customer environments: mixed workloads (block, file, object), long‑running IO, failovers, NDUs, hardware events, and background operations (replication, snapshots, quotas, etc.), combining automation with targeted “tortures”. Drive execution and triage on System Testing beds Own System Testing environments (arrays, initiators, OSes, accessories); keep them healthy, representative, and well‑instrumented. Monitor runs, triage failures quickly, separate infra issues from product bugs, and file high‑quality defe

pythonawskubernetes
View job →
TA
22 days ago

About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw

pythonkuberneteslinux
View job →
G
Glance
📍 Bengaluru• Full-time
22 days ago

Glance AI is an AI commerce platform shaping the next wave of e-commerce with inspiration-led shopping, less about searching for what you want and more about discovering who you could be. Operating in 140 countries, Glance AI transforms every screen into a stage for instant, personal, and joyful discovery, where inspiration becomes something you can explore, feel, and shop in the moment. Its proprietary models, seamlessly integrated with Google’s most advanced AI platforms, Gemini and Imagen on Vertex AI, deliver hyper-realistic, deeply personal shopping experiences across categories such as fashion, beauty, travel, accessories, home décor, pets, and more. Designed to seamlessly integrate into everyday consumer technology, Glance AI reimagines the future of e-commerce with inspiration-led discovery and shopping. With an open architecture built for effortless adoption across hardware and software ecosystems, Glance AI is creating a platform that can become a staple in everyday consumer technology. It partners with the world’s leading smartphone makers, connected TV manufacturers, telecom providers, and global brands — meeting people where they are: on mobile, smart TVs, and brand websites. Through Glance AI’s rich first-party data and unparalleled consumer access, it harnesses InMobi’s global scale, insights, and targeting capabilities to create high-impact, performance-driven shopping journeys for brands worldwide. Part of the InMobi Group, a global technology and advertising leader reaching over 2 billion devices and serving more than 30,000 enterprise brands worldwide, Glance AI is backed by Google, Jio Platforms, and Mithril Capital. Strategy & Operations Manager – Glance TV Experience: 5 to 7 years We are looking for a highly analytical and execution-oriented Strategy & Operations Lead to support the growth of Glance TV across markets. This role will work closely with Product, Engineering, Design, Content, Ads Monetization, Finance, and Analytics

M
Modal
📍 New York• Full-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track

aigorust
View job →
N
Nuro
📍 Mountain View• Full-time• From $160.4K/yr
1mo ago

Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Team The Systems Engineering team is responsible for the requirements, architecture, and validation of autonomous driving capabilities across engineering disciplines. This includes designing performance metrics, evaluation methods, and criteria for success, which the team then drives cross-functional via requirement definition and system validation. Systems Engineering works at the intersection of hardware, software, and robot operations, with a deep understanding of technologies in all three. We are a small, high-impact team that sets the checkpoints for autonomy deployment. About the Role The Senior Systems Test Engineer (STE), Autonomy Behavior role is responsible for transforming Nuro's Verification & Validation (V&V) ecosystem. You will focus on the technical implementation, standardization, and automation of V&V processes. This involves creating a unified, automated architecture for validation , designing and

pythonaic++
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $96.3K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As an IT Service Desk Engineer on our IT team, you'll play a key role in keeping Coinbase employees productive and supported globally. You'll own complex Tier 2 escalations, solve issues across devices, applications, and cloud platforms, and help improve and scale the service desk as Coinbase grows. What you'll do: Own Tier 2 troubleshooting and resolution for hardware, software, networking, and access issues escalated from Tier 1 while meeting SLAs Support Coinbase's SaaS ecosystem including Google Workspace, Atlassian, Okta, Jamf, and Slack, along with compliance and onboarding/offboarding workflows Perform root cause analysis for recurring or high-impact issues and drive corrective actions through resolution Build and maintain documentation, SOPs, and knowledge base content that reduce repeat tickets and improve support quality Partner with IT Operations, Security, IAM, Corp Engineering, Workplace, Procurement, and vendors to resolve advanced issues and improve service delivery Identify support trends and workflow gaps, then use automation, scripting, or process improvements to reduce manual work and deliver measurable operational impact Required Skills and Experience: 5+ years of Tier 2 technical support experience across macOS, iOS, Chrome OS, Android, and Windows/AWS Workspaces, including 3+ years administering Google Workspace, Atlassian, Okta, Jamf, and Slack

REMOTEawsagileai
View job →

At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Lyft connects people to transportation to change the way we live and get around our communities. Lyft Urban Solutions (LUS) is North America’s leader in micromobility. We own, operate, or provide hardware and software solutions for bikeshare and scootershare programs in 57+ global markets including Montreal, London, New York City, Mexico City and others. Our rapidly growing active fleet includes state-of-the-art charging stations, electric bikes, and scooters, and services hundreds of millions of rides per year. The Hardware Product and Industrial Design Team at LUS is responsible for conceiving, guiding development of, launching, and sustaining world-class products used by riders, cities, and operators over the world. We represent LUS in micromobility forums, engage directly with customers, and guide research and development. We aim to ensure all products we develop are delightful for riders, globally competitive, elegant, and future-facing. We are looking for an exceptional Senior Product Manager, Hardware to manage our existing stations portfolio as well as develop our next generation of stations products. Our stations are now as iconic in cityscapes as London Routemaster buses of old. You will have end-to-end ownership of new station product development, leading a cross-functional team through all phases of hardware development to launch, and will be responsible for a fleet of over 300,000 charging and non-charging station assets already on the ground. You will engage directly with LUS leadership, engineering, data science, operations, and design to achieve the company’s vision for reinventing transportation. You will work closely with our Vehicles Product Manager, Industrial Design team, and counterparts in software product management to ensure ou

S
Stripe
📍 Nyc Privy• Full-time
1mo ago

Who we are About Privy Our mission is to make privacy and user ownership the default online. To do so, we build simple, flexible APIs and tools for developers that make it easy to build new products on crypto rails. Privy owns the abstractions and infrastructure layer above wallets, integrating across chains, third-party providers, and Stripe products like Treasury and Link. We get to solve hard technical problems while leveraging Stripe's distribution to reach customers like Ramp, Klarna, Deel, Kraken, Hyperliquid, and Fomo — powering experiences for both mainstream users and crypto natives. Learn more about Privy: Privy and Stripe: Bringing crypto to everyone About the team Engineering at Privy is distinguished by: High urgency: Shipping very small iterations, very fast, to learn very quickly. Product taste: Our customers are developers, and to build effective products for them requires technical knowledge - you will often be "the PM". Security mindset: A great portion of our product is trust. While we have a dedicated security team, every engineer brings security to their designs from the start. In practice, we use boring technology like Node, React, and AWS so we can focus our engineering energy entirely on pushing the boundaries of Privy's core product, e.g. through hardware enclaves, multi-region low latency APIs, and blockchain abstractions that are accessible to mainstream developers. What you’ll do As a Security Engineer at Privy, you will help keep a rapidly growing platform secure through hands-on investigation, operational ownership, and building real code to solve real problems. You’ll work across security operations, incident response, application and infrastructure security, and internal tooling. This is not a role limited to reviewing logs and designs or simply escalating findings. You will carry real ownership for the day-to-day work that protects Privy: investigating anomalies, improving security workflows, and building the automa

pythonreactaws
View job →

About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man

pythonsqlaws
View job →
O
1mo ago

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

pythonawslinux
View job →
O
1mo ago

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

awsrestai
View job →
C-
CLEAR - Corporate
📍 New York• Full-time• From $145K/yr
22 days ago

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a Senior Manager of Industrial Engineering to join CLEAR's Central Operations team. This role is the quantitative backbone of our Experience & Operations Design function — owning the data-driven design standards, labor models, and capacity planning that determine how CLEAR's physical operation is structured and staffed. You will be the analytical engine behind lane design decisions, engineered labor standards, and demand forecasts, ensuring the operation is designed to deliver a consistently excellent member experience efficiently at scale. What you'll do: Develop and own engineered labor standards and staffing models across CLEAR’s physical experiences, including Verifier, Greeter, eGate, Concierge, TSA PreCheck, Enrollment, Sports, and new programs, serving as the operations source of truth for labor and headcount inputs Define physical operating requirements and sizing standards across lane configurations, verification and enrollment experiences, hardware placement, and queue management; lead checkpoint design for new and existing airport locations Lead demand forecasting and capacity planning across CLEAR experiences, modeling throughput, staffing requirements, operating hours, and operational impact for new programs, launches, surge periods, and events Drive operational efficiency through time studies, process analysis, and labor optimization, identifying and developing opportunities to improve productivity, capacity, and the Member experience across the network Partner cross-functionally with Operations,

pythonsqlgit
View job →
S
Stripe
📍 San Francisco• Full-time
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team Stripe Terminal helps businesses extend their online presence into the physical world. Our mission is to make it as easy for businesses to accept in-person payments as Stripe has made it to accept payments online. Terminal provides an integrated in-person payments platform spanning software, payment processing, device fleet management, and hardware. Our hardware portfolio includes Stripe-designed readers and accessories as well as devices from third-party hardware partners. Together, these products enable businesses to build reliable, differentiated in-person payment experiences across a broad range of use cases, countries, and operating environments. What you’ll do We’re looking for an experienced Product Manager to build the strategy, development, and expansion of Terminal’s hardware portfolio. This role will work across both first-party and third-party hardware, helping define the multi-year roadmap for the devices, accessories, and partner ecosystem that power Terminal. You will partner closely with hardware engineering, firmware and software engineering, design, operations, logistics, partnerships, sales, finance, legal, and regional teams. This is a highly cross-functional role for someone who can connect customer needs, technical constraints, business strategy, and execution across user experience, device capabilities, cost, quality, reliability, supply, and time to market. Responsibilities Set and execute a multi-year strategy and ro

aigofinance
View job →
H
Hyliion
📍 Austin• Full-time
22 days ago

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Technical Program Manager owns the planning and coordination of hardware programs focused on our assembling and other traditionally manufactured power generation components and deliverables. This role helps bridge the gap between engineering, supply chain, manufacturing, and operations, ensuring our beta-stage designs make a successful leap to production. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Drive the end-to-end program execution of assembling the components, from prototype validation to initial production ramp. Partner with mechanical design engineers and manufacturing engineers to ensure designs are manufacturable and scalable. Develop and maintain detailed project timelines, tracking milestones, budgets, and risk mitigation plans. Strategic Problem-Solving: Apply a structured, hypothesis-driven approach to diagnose complex program challenges, identify root causes, and develop data-backed solutions. Lead cross-functional meetings to align engineering, operations, procurement, quality assurance and external vendors around shared goals. Coordinate testing and validation of parts, including mechanical, thermal, and durability testing in real-world power generation applications. Assist in supplier sourcing, DFM (Design for Manufacturability) reviews, and process validation with supplier partners. Track any ECRs resulting from modifications needed from design issues, testing, DFM or DFMEA ac

aigoexcel
View job →
H
Hark
📍 San Jose• $200K – $400K/yr
15 days ago

About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory. We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world. To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems. About the Role You own Hark's hardware distribution metric. That means the strategy and the execution behind getting our device into the market at scale in its first months. You'll be the central quarterback across carriers, engineering, product, and the business teams, holding the single Concept of Operations that everyone works from and rewriting it when the data says the plan isn't working. This is a high-ownership role where the target is measurable and yours. Responsibilities Distribution Target Ownership: Carry the unit sales number end to end, and own the plan that gets Hark there on the launch timeline. Concept of Operations: Build and maintain the single central ConOps that defines how distribution runs — channels, sequencing, dependencies, and who owns what. Channel Strategy and Execution: Turn carrier and retail relationships into committed sell-in and real sell-through, from first conversation to units moving. Cross-Functional Alignment: Pull input from carriers, engineering, product, and business teams into one plan, and keep those teams working against the same commitments. Performance Review and Adaptation: Run a constant loop on data, results, and channel feedback, and change the plan the moment the numbers say it isn't working. Launch R

artificial intelligenceai
View job →
🔔

Get new hardware operations engineer jobs by email

Daily job updates · Unsubscribe anytime