Jobiba hiring network

Hardware Operations Engineer Jobs

1,283 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current hardware operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
1mo ago

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but operational in how we execute, and we support every product and research effort at OpenAI. Our tenets include prioritizing for impact, enabling researchers and developers, preparing for future transformative technologies, and fostering a strong, collaborative security culture. About the Role OpenAI is seeking a Principal Software Engineer to join the Infrastructure Security (InfraSec) team. InfraSec safeguards the core of OpenAI’s research and production environments: GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter spans everything from bare-metal hardware and firmware to Kubernetes clusters, service meshes, and the data pathways that carry highly sensitive model weights and user data. As a Principal Software Engineer, you will set technical direction and drive execution of critical foundational services, such as authentication systems, egress/ingress proxies, access brokers, and key management platforms, that demand high standards of reliability, scalability, and software craftsmanship. These systems form the security backbone of OpenAI’s customer and supercomputing environment and must remain robust under intense scale and adversarial pressure. In this role, you will: Own the architecture and roadmap for one or more core security services (e.g., authN/Z, policy enforcement, secure proxies, key management), taking them from design to rollout to long-term operation. Design and implement planet-scale security systems that provide strong guarantees across hardware, operating systems, Kubernetes, networks, and CI/CD: balancing security, reliability, latency, and developer ergonomics. Lead cross-functional launches

awsazuregcp
View job →

Radiation Environmental Test Engineer (Mid-Level or Senior) Company: The Boeing Company Boeing Missile & Weapons Systems is seeking a Radiation Environmental Test Engineer (Mid-Level or Senior) to lead radiation and environmental test efforts supporting ICBM modernization at the Little Mountain Test Facility in Hill Air Force Base, UT. This hands-on role combines test design, execution, data analysis, and facility/test-stand maintenance. The position supports ongoing modernization efforts to enhance current test capabilities and new developing capabilities—identifying, designing, managing, and integration of new test stands, diagnostics, and data systems to meet evolving mission requirements. Position Responsibilities: Lead and execute complex radiation and environmental test machine design and facility integration Identify interface requirements, CDRLs, configuration control documentation, and change orders Design, specify, and integrate test hardware, sub-assemblies, fixtures, and instrumentation Assist in configuring, validating, and operating data acquisition systems and custom software (LabVIEW) for diagnostics and event capture Support pulse power and flash X-ray systems: theory, operation, diagnostics, and data capture Perform technical evaluation of subcontractors Basic Qualifications (Required Skills and Experience): Bachelor of Science degree in Engineering (with a focus in Electrical, Mechanical or Aeronautical), Computer Science, Data Science, Mathematics, Physics, Chemistry or non-US equivalent qualifications directly related to the work statement Level 3: 5+ years of related work experience and a bachelor’s degree or an equivalent combination of related work

recruitment
View job →

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Join Our Team as a STAFF ENGINEER, TDT PD CMP At Micron, we are driven by the pursuit of excellence and innovation. As a Chemical Mechanical Planarization (CMP) Process Development Engineer, you will play a key role on our Technology Development Taiwan (TDT) CMP team. You will contribute to the progress of our innovative DRAM technology. If you have a passion for semiconductor development and enjoy working in a dynamic and collaborative environment, we invite you to join us in defining the future of the semiconductor industry! Key Responsibilities Develop and refine CMP processes to support advanced DRAM technology acceleration. Collaborate with diverse teams to develop new outstanding semiconductor devices. Perform rigorous process characterization, develop new consumables, and evaluate new hardware. Conduct groundbreaking research and innovate to contribute to Micron’s extensive patent portfolio. Drive continuous process and efficiency improvements to meet yield, quality, and cycle time requirements. Direct and manage CMP consumables and equipment suppliers. Plan and complete well-designed experiments, presenting results in clear and concise reports. Minimum Qualifications Proficiency in both Chinese and English languages. Expertise in CMP process development. Hands-on experience with CMP polishing equipment, including familiarity with equipment operation, optical metrology, profilometry, common polishing

machine learningartificial intelligenceai
View job →
B
18 days ago

Human Engineer (Associate or Experienced) Company: The Boeing Company Are you ready to join a team of innovators, strategic disruptors, and dreamers who dare to redefine the future of the defense industry? Are you driven to create the unimaginable? Are you passionate about addressing human performance within the design and development of cutting-edge aerospace systems? Boeing Defense, Space & Security (BDS ) has an exciting opportunity for an associate or experienced Human Engineer to join our Human Engineering team in Colorado Spings, CO working on the front end of the next generation of Space Mission Systems technology. As a Human Factors Engineer, you will play a pivotal role in the design, development, and integration of complex aerospace systems, ensuring that all components function cohesively to meet user needs and operational requirements. You will collaborate with cross-functional teams, including systems engineers, software developers, and hardware specialists, to create systems that are not only technically robust but also intuitive and user-friendly, enhancing overall efficiency and user satisfaction. Position Responsibilities : Apply Human Engineering knowledge and principles to the analysis, design, and evaluation of complex systems Define system performance requirements to ensure safe and successful human operation of physical, functional, and program interfaces for all system operators Apply human performance principles, methodologies, and technologies to the design of complex systems Develop and implement research methodologies and analysis plans to test and evaluate developmental prototypes Apply a knowledge of military standards (such as MIL-STD-1472, MIL-STD-1474,

recruitment
View job →

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE This role is for an individual contributor joining one of our five Everpure Resilience engineering teams within the Data & Digital Experience (DDX) organization. The product has been running for 3 years and operates like a startup within DDX. The teams are primarily composed of senior-level engineers who take full ownership of their workstreams - from early requirements definition (in close collaboration with Product Managers) to final delivery. Strong communication and collaboration skills are essential, as engineers work autonomously in a highly cross-functional environment. WHAT YOU'LL DO Architect & Deliver: Own the end-to-end design, development, and operation of high-throughput data processing services between edge devices and the Everpure cloud platform. Product Collaboration: Partner with Product Managers to translate complex requirements into scalable, resilient system architecture from concept to production. Performance & Scale: Drive continuous system improvements—experimenting with new tech to boost performance, security, and cost-efficiency for real-time data flows. Ship Fast & Safely: Utilize pragmatic testing, robust CI/CD, and automation to maintain high deployment velocity without sacrificing stability. System Integration: Lead the resolution of complex interoperability challenges between new and legacy components to ensure high availability across all regions. WHAT YOU BRING Distr

javaawsazure
View job →
O
1mo ago

About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo

awslinuxrest
View job →

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across silicon, embedded systems, operating systems, and cloud services to build reliable consumer devices and the novel platforms required to support them. We partner closely with research to bring advanced AI capabilities into the physical world. About the Role As an Operating Systems Engineer focused on on-device inference, you will design, develop, and ship the OS stack that makes advanced AI capabilities reliable, responsive, and energy efficient on consumer devices. Your work will span OS services and frameworks, inference runtime integration, model fitting, scheduling, and performance and power management. You’ll partner with research to adapt models to device constraints, make design decisions across the stack, and carry solutions from early exploration through integration and production. In this role, you will: Build the inference platform: Design and implement maintainable OS services, frameworks, and clear interfaces for inference execution, model loading and lifecycle, and resource management. Fit models to device constraints: Partner with researchers on quantization, runtime integration, and memory optimization to meet memory, compute, and energy budgets while evaluating model quality and product behavior. Coordinate system resources: Develop scheduling and resource policies that balance inference with other device act

machine learningartificial intelligenceai
View job →

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across custom silicon, embedded systems, operating systems, and cloud services to build reliable consumer devices and the platforms behind them. We connect kernel development with the broader software stack to deliver complete product capabilities. About the Role As an Operating Systems Engineer focused on the Linux kernel, you will design, develop, and maintain the kernel capabilities that underpin OpenAI’s consumer devices. You’ll bring deep expertise in one or more Linux kernel subsystems and carry solutions through the higher-level software stack. Your ownership will extend into the userspace services, libraries, tools, and interfaces needed to deliver complete product features. You’ll shape the boundaries between kernel and userspace, make design decisions across the stack, and see your work through development, integration, and production. In this role, you will: Build kernel capabilities: Design, implement, and maintain Linux kernel subsystem changes that support device capabilities and product requirements. Own features across the stack: Choose appropriate kernel and userspace boundaries, and build the interfaces and supporting components needed to deliver reliable features in shipped products. Debug complex system behavior: Use tracing, profiling, instrumentation, and diagnostic tools to resolve correctness, concurrency, p

linuxartificial intelligenceai
View job →

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across systems software and product engineering to build reliable consumer devices and the platforms behind them. We develop the connectivity and networking foundations that support communication across OpenAI products and systems. About the Role As an Operating Systems Engineer focused on connectivity and networking, you will design, develop, and maintain the OS capabilities that enable reliable, secure, and efficient communication across OpenAI products and systems. Your work will span Wi-Fi and Bluetooth frameworks, IP networking, and advanced network services and policy. You’ll develop OS services, libraries, and interfaces for a broad range of connectivity needs, make design decisions across software boundaries, and carry solutions through development, integration, and production. In this role, you will: Build connectivity foundations: Design, implement, and maintain OS services, frameworks, and APIs for Wi-Fi, Bluetooth, and IP networking. Develop reusable network capabilities: Build connection management, network configuration and selection, service discovery, and routing capabilities. Enable secure communication: Develop network services and policies for secure communication, traffic management, and isolation across varied network environments. Resolve issues across the stack: Investigate correctness, concurrency, interoper

artificial intelligenceaic++
View job →
O
1mo ago

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role summary We are seeking a Networking Operating System Firmware Engineer to help bootstrap and scale the switching layer of our AI supercomputers. In this role, you will build and maintain custom NOS images from scratch, using open source components from SONiC, SAI, FRR, and related networking stacks while working across the Linux kernel, switch ASIC SAI/SDKs, platform drivers, control-plane services, and orchestration layers. This is a software engineering role that requires a deep understanding of networking, NOS internals, switch hardware, and production systems. You will design, implement, test, and debug production NOS software across platform drivers, routing and control-plane state, ASIC programming, observability, and fleet integration. The engineer in this role should be able to work through ambiguous, open-ended technical problems and drive feature development across software, hardware, and vendor boundaries. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Design, develop, and maintain custom NOS images for large-scale AI fabrics, using open source components from SONiC, FRR, and related networking stacks. Integrate, build and configure Linux kernel components, device drivers, switch ASIC SDKs, and SAI layers. Bring up new switch platforms, including thermal and fan control, power monitoring, transceiver management, watchdogs, OSFP CMIS, L

pythonawsci/cd
View job →
E
23 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pipeline Engineering team at Everpure™ as a Software Engineer to build, own, and operationally scale the microservices and Temporal workflows driving continuous integration for FlashArray, FlashBlade, and Hyperscale. In this developer-first role based in Bangalore, you will design durable production services and agentic AI tools that directly reduce time-to-signal, optimize compute utilization, and absorb operational toil across our global engineering workforce. WHAT YOU'LL DO Architect Deterministic CI Workflows: Design and deploy production microservices and durable Temporal workflows that make pipeline execution resilient, resumable, and fully debuggable for enterprise storage platforms. Optimize Compute & Testbed Utilization: Extend in-house scheduling logic and placement algorithms across bare-metal hardware and VM fleets to minimize queue times and maximize infrastructure efficiency. Build Operational AI Agents: Engineer RAG pipelines and agentic workflows over failure logs, vector stores, and test metadata to automate root-cause analysis, flake classification, and developer triage. Drive Platform Observability & Ownership: Establish key platform metrics—including time-to-signal and pass/flake rates—while operating your services end-to-end to ensure long-term stability and eliminate recurring failure modes. Collaborate Across Product Teams: Partner with platform and product engineering group

pythonsqlaws
View job →
E
Everpure
📍 Bengaluru• Full-time
23 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Pipeline Engineering team at Everpure™ as a Software Engineer to build, own, and operationally scale the microservices and Temporal workflows driving continuous integration for FlashArray, FlashBlade, and Hyperscale. In this developer-first role based in Bangalore, you will design durable production services and agentic AI tools that directly reduce time-to-signal, optimize compute utilization, and absorb operational toil across our global engineering workforce. WHAT YOU'LL DO Architect Deterministic CI Workflows: Design and deploy production microservices and durable Temporal workflows that make pipeline execution resilient, resumable, and fully debuggable for enterprise storage platforms. Optimize Compute & Testbed Utilization: Extend in-house scheduling logic and placement algorithms across bare-metal hardware and VM fleets to minimize queue times and maximize infrastructure efficiency. Build Operational AI Agents: Engineer RAG pipelines and agentic workflows over failure logs, vector stores, and test metadata to automate root-cause analysis, flake classification, and developer triage. Drive Platform Observability & Ownership: Establish key platform metrics—including time-to-signal and pass/flake rates—while operating your services end-to-end to ensure long-term stability and eliminate recurring failure modes. Collaborate Across Product Teams: Partner with platform and product engineering group

pythonsqlaws
View job →
G
23 days ago

1743 - This position is in Austin, Texas. Position Summary We are seeking an experienced Board-Level Hardware Validation Engineer to define and execute the validation and verification of complex electronic systems throughout the product lifecycle. This role is responsible for defining validation strategies, developing test plans, executing hands-on testing, analyzing failures, and working directly with ODM partners to ensure products meet performance, reliability, quality, and compliance requirements before mass production. The ideal candidate combines strong electrical engineering fundamentals with practical lab expertise and is comfortable personally performing validation activities while coordinating with cross-functional teams and manufacturing partners. Key Responsibilities Validation Strategy & Planning Define comprehensive board-level and inter-board validation plans based on product requirements, design specifications, and customer use cases. Develop validation methodologies covering functional, electrical, thermal, power, signal integrity, reliability, and stress testing. Establish test coverage, acceptance criteria, qualification requirements, and release gates. Review hardware architecture, schematics, component specifications, and interface topologies to identify validation risks early in the design cycle. Define incremental validation and regression coverage for component substitutions, design changes, and firmware updates. Hands-On Validation Execution Develop, automate, and execute validation tests on prototype and production-intent hardware. Perform board bring-up, functional verification, electrical characterization, and system-level integration testing. Validate communication interfaces, control signals, and timing requirements. Verify power sequencing, reset behavior, leakage current, and recovery across operating states. Execute temperature and voltage corner testing against approved operating limits. Use oscilloscopes, logic an

pythongitai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

pythonsqlaws
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work

pythonsqlaws
View job →
🔔

Get new hardware operations engineer jobs by email

Daily job updates · Unsubscribe anytime