Jobiba hiring network

Hardware Operations Engineer Jobs

1,283 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current hardware operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →
O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →
S
SingleStore-LinkedIn
📍 San Francisco• Full-time• $960K – $1.4M/yr
21 days ago

Position Overview As SingleStore’s IT Operations Engineer, you will help shape the IT toolset used by our end users. This is an active, hands-on position responsible for the planning, design, development, and Tier 1 support of several key technical areas at the SingleStore IT team, including end-user support, client engineering, executive support, and infrastructure application support. This is an incredible opportunity for someone to build upon their technical strengths and be a part of IT at SingleStore team . Roles and Responsibilities: Administering a wide variety of SaaS applications. Some main applications that need to be supported are OKTA (+ Workflows), Google Workspace, Slack, and Atlassian tools (JIRA + Confluence), MDM administration. Keep up to date with new features and new releases in these applications to identify opportunities for better automation or features that could be useful for our environment. Seize opportunities across the IT Operations team to eliminate manual work through tooling, integrations, and automation of IT workflows. Respond to tickets and execute new hire onboarding and user separation processes. Support members of the team with troubleshooting and resolution of complex issues. Design, architect, implement and maintain systems and solutions for various IT-related topics, including but not limited to staff computer hardware, operating systems, software applications, networking, videoconferencing, and printers. Partner and collaborate with all business units to help them evaluate hardware and software solutions. Able to communicate effectively and concisely with the entire company. Analyze existing processes, suggest and make improvements, and implement business processes where none exists. A desire to learn and expand your horizons; take on new challenges as the business scales Required Skills and Experience: Minimum 2 years of relevant experience Prior experience in implementing and administering Google Workspac

pythonsqlaws
View job →
E
21 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a Staff Security Operations Engineer , you will own and continuously mature capabilities across Attack Surface Management, Vulnerability Management, Zero Trust, Secrets and Credential Security, Detection Engineering, and Security Automation . This is a hands-on role requiring strong security engineering expertise combined with the ability to build teams, establish operating processes, measure outcomes, and drive remediation across Engineering, Cloud, Infrastructure, and Product organizations. You will partner closely with GISO leadership and global security teams to translate security strategy into measurable execution and risk reduction. WHAT YOU’LL DO Lead and mature enterprise Attack Surface and Vulnerability Management capabilities across cloud, infrastructure, endpoints, applications, and internet-facing environments. Drive risk-based vulnerability prioritization using asset criticality, exposure, exploitability, known exploitation, threat intelligence, and compensating controls. Establish operating processes, remediation SLAs, KPIs/KRIs, dashboards, and governance to measure and drive security risk reduction. Identify systemic security gaps and develop scalable technical and operational solutions. Provide technical leadership across Zscaler/Zero Trust, secrets and credential security, SIEM/detection engineering, EDR, cloud security, and security automation. Drive automation and integrations using APIs, sc

awsazuregcp
View job →
E
21 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a Senior Security Operations Engineer – Attack Surface Management , you will help engineer, operate, and continuously improve our enterprise Attack Surface and Vulnerability Management capabilities. You will focus on discovering and understanding our attack surface, identifying vulnerabilities and exposures, prioritizing them based on real-world risk, and driving remediation across our global environment. The role spans Attack Surface Management, Vulnerability Management, Zero Trust, Secrets and Credential Security, SSPM, Deception, Detection Engineering and Security Automation. We are looking for engineers with strong ASM/Vulnerability Management fundamentals and deeper expertise in one or more adjacent security domains. WHAT YOU’LL DO Engineer and improve enterprise Attack Surface and Vulnerability Management across cloud, infrastructure, endpoints, applications, and internet-facing environments. Discover and correlate assets across security platforms and identify unknown, unmanaged, stale, and externally exposed assets. Drive risk-based vulnerability prioritization using asset criticality, exposure, exploitability, known exploitation, threat intelligence, and compensating controls. Establish remediation workflows and SLAs and partner with asset owners to drive measurable risk reduction. Operate and troubleshoot Zscaler ZIA/ZPA, including SSL inspection, access policies, connectivity, and Zero Trust controls.

pythonawsazure
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set

machine learningaigo
View job →
O
OpenAI
📍 Singapore• Full-time
1mo ago

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. We work closely with hardware design teams, manufacturing partners, suppliers, and data center operations to deliver the compute platforms that power frontier AI. The Manufacturing Engineering team ensures our hardware can be built, tested, deployed, and supported at hyperscale. We bridge Hardware Engineering, Quality, Supply Chain, Contract Manufacturers, and Hardware Operations to continuously improve manufacturing performance throughout the product lifecycle. As our fleet grows globally, sustaining manufacturing engineering becomes increasingly important to maintain product quality, improve manufacturability, and rapidly resolve production issues. About the Role We are seeking a PCBA Manufacturing Engineer (Sustaining) to support production and continuous improvement of printed circuit board assemblies (PCBAs) used throughout Industrial Compute hardware platforms. This role focuses on sustaining engineering after product launch. You'll partner closely with Hardware Design, Quality, Manufacturing, Test Engineering, Supply Chain, and our contract manufacturers to resolve production issues, improve manufacturing yield, reduce failures, implement engineering changes, and ensure stable high-volume manufacturing. The ideal candidate has experience supporting complex server, networking, accelerator, storage, or high-performance electronics manufacturing environments. Key Responsibilities Own sustaining manufacturing engineering for PCBA production across multiple hardware platforms. Drive root cause investigations for manufacturing defects, field failures, and production escapes. Partner with Hardware Design Engineers to improve manufacturability (DFM/DFA/DFT). Support engineering change orders (ECOs) and manufacturing change implementation. Work directly with contract manufacturers to improve production yield, cycle time, and quality. Analyze manuf

awsrestai
View job →
O
OpenAI
📍 Singapore• Full-time
1mo ago

About the Team The Consumer Products team at OpenAI brings groundbreaking AI technology to life through world-class hardware. Our Operations group ensures that every product moves seamlessly from concept to customer — integrating supply chain strategy, manufacturing operations, and quality systems to enable rapid, reliable, and scalable production. We partner closely with engineering, design, and manufacturing teams around the world to deliver exceptional products at launch and beyond. About the Role As an Operations Manager, you will own the execution and strategy that transforms early product concepts into high-quality, manufacturable consumer devices. You’ll lead cross-functional planning, manage vendor relationships, and ensure supply, cost, and schedule alignment across fast-moving hardware programs. This role is based in Singapore. We follow a hybrid model (four days per week in the office) and offer relocation support for new employees. Occasional travel to supplier and manufacturing sites is expected. In this role, you will: Develop and execute operations strategies that bridge product design and large-scale manufacturing. Drive build readiness and production planning from prototype through mass production. Partner with global suppliers to ensure materials, tooling, and capacity meet program goals. Collaborate with hardware, quality, and supply chain teams to establish and maintain robust operational systems. Track cost, yield, and schedule performance to ensure efficient and predictable product delivery. Build scalable processes that support rapid iteration and future product launches. You might thrive in this role if you: Have 6+ years of experience in hardware operations, manufacturing, or supply chain for consumer electronics or similar industries. Have successfully taken at least one product from concept through mass production. Are fluent in NPI processes, contract manufacturing dynamics, and global operations management. Bring strong analytical and co

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o

awsrestai
View job →
O
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world's most advanced AI infrastructure ecosystem. The Scaling Analytics team serves as the data backbone for this effort, enabling leaders and operators to make informed decisions across infrastructure deployment, hardware operations, supply chain, capacity planning, and site execution. As OpenAI’s Industrial Compute expands across an increasing number of global data center campuses, the complexity of managing infrastructure capacity, hardware health, supply flows, and operational performance continues to grow. Scaling Analytics develops the data models, pipelines, metrics, and reporting systems that transform fragmented operational data into actionable insights, helping OpenAI operate infrastructure at unprecedented scale. About the Role We are seeking a Data Engineer to help build and scale the analytical foundations that power OpenAI's infrastructure organization. This individual will partner closely with Hardware Operations, Capacity Planning, Supply Chain, Infrastructure Delivery, Finance, and Engineering teams to create reliable data products that support critical operational and strategic decisions. Today, much of the team's expertise is concentrated within several highly specialized domains including hardware health, GPU attribution, and supply analytics. As Stargate grows and new sites come online, the demand for analytics support continues to expand across both existing and emerging problem spaces. This role will increase the team's ability to move quickly, reduce operational bottlenecks, and provide additional depth across critical infrastructure analytics functions. The ideal candidate combines strong data engineering fundamentals with an ability to navigate ambiguous operational environments, translating complex infrastructure problems into scalable data solutions that improve visibility, decision-making, and execution. Key Responsibilities Design, build, and maint

pythonsqlaws
View job →
O
29 days ago

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to develop and operate increasingly capable AI systems at global scale. The Manufacturing Operations team works across Hardware Engineering, Manufacturing Engineering, Manufacturing Quality, Rack Integration, System Enablement, Supply Chain, Logistics, Deployment, and Hardware Operations to convert complex hardware designs into reliable production systems. We partner closely with ODMs, JDMs, contract manufacturers, and component suppliers to ensure that servers, racks, networking equipment, and supporting infrastructure are manufactured, validated, and delivered at the quality and scale required by OpenAI. About the Role We are seeking a Technical Program Manager to lead manufacturing operations programs across OpenAI’s hardware supply base in Singapore and the broader APAC region. You will own cross-functional execution from new product introduction through production ramp and sustaining operations. You will coordinate manufacturing partners and internal engineering teams around factory readiness, capacity, material availability, build plans, validation, quality gates, issue resolution, and delivery commitments. This role requires strong technical fluency, disciplined program management, and the ability to operate directly with manufacturing partners in fast-moving, high-stakes environments. You should be comfortable working at both the factory floor and executive-review levels, translating complex manufacturing conditions into clear risks, decisions, and recovery plans. This role is based in Singapore and requires regular travel to manufacturing partners across the APAC region. Key Responsibilities Lead manufacturing operations programs for AI servers, racks, networking systems, and related infrastructure across regional manufacturing partners. Own integrated program plans spanning NPI, factory readiness, material availability, tooling, test development, qualification,

REMOTEpythonsqlaws
View job →

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. The team ensures that large, highly complex infrastructure programs execute predictably across planning, design, construction, hardware deployment, commissioning, and operational handoff. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Supply Chain, and our external infrastructure partners to keep programs aligned, risks visible, and execution moving at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive execution across large-scale AI infrastructure deployments. This role is responsible for orchestrating cross-functional delivery programs spanning multiple organizations, ensuring dependencies remain synchronized from early planning through production readiness. You will develop operational mechanisms that allow Industrial Compute to scale infrastructure delivery across multiple campuses simultaneously. Rather than owning any individual engineering discipline, you will own program health—bringing together engineering, construction, operations, supply chain, and partner organizations into a single coordinated execution model. Success in this role requires exceptional program management, executive communication, systems thinking, and operational rigor. You should be comfortable operating amid ambiguity while bringing structure to highly technical, multi-year infrastructure programs. Candidates should have experience leading complex infrastructure, cloud, data center, semiconductor, networking, manuf

awsrestagile
View job →

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. Our team develops the operating model that connects infrastructure strategy, supply planning, manufacturing operations, and delivery into a single, integrated system that enables OpenAI to deploy AI infrastructure predictably at scale. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Strategic Sourcing, and external infrastructure partners to create a single, integrated view of program health. Through governance, operational analytics, executive reporting, and scalable operating mechanisms, we enable leaders to proactively manage risk, optimize capacity, and deliver infrastructure predictably at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive integrated strategy and delivery across OpenAI's rapidly expanding AI infrastructure portfolio. This role sits at the intersection of infrastructure strategy, New Product Introduction (NPI), supply planning, manufacturing operations, and infrastructure delivery. You will lead highly cross-functional programs spanning engineering, supply planning, manufacturing, logistics, construction, commissioning, and operations, ensuring technical and operational dependencies remain synchronized from planning through production readiness. Beyond driving program execution, you will leverage operational insights to improve capacity planning, infrastructure strategy, and deployment readiness. You will also help operationalize new technologies and suppliers by partnering w

REMOTEawsrestagile
View job →
O
1mo ago

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We build the shared infrastructure, tooling, and lab environments that let teams test quickly, understand failures, and ship with confidence. About the Role As a Systems Integration Manager , you will lead the team responsible for device validation infrastructure, test automation, developer tooling, and lab operations. This is a player-coach leadership role: you’ll set technical and operational direction, build and develop a team of engineers and lab operations professionals, and stay close to the architecture and hardest systems problems. You will partner closely with device software, OS, firmware, hardware, reliability, QA, and release infrastructure teams to define validation strategy, improve release readiness, and ensure our test environments and quality signals scale with the product. Because this is a new category of devices, you’ll have the rare opportunity to build the validation foundation early—shaping the systems, standards, and operating model that will support products from prototype through launch. We’re looking for a leader who combines strong technical judgment with people leadership, operational rigor, and experience building reliable systems for complex hardware-software products. This role is b

artificial intelligenceai
View job →

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure: Owner Furnished Equipment to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastr

awsrestai
View job →
🔔

Get new hardware operations engineer jobs by email

Daily job updates · Unsubscribe anytime