About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
Site Reliability Engineer / Production Support Analyst
Role market pulse
How DevOps Engineer demand looks in United States
Live jobs
17
Posted 30d
4
30d movement
-69.2%
Remote share
17.6%
Salary listed
58.8%
Salary trend 1Y
Not enough history
Role overview
Job description
Location: South Jordan, UT (This role is on-site) Salary: $56,000 - $58,000 USD Want to work in technology at an investment bank? Graduate training, ongoing support, opportunities at leading global employers – the Alumni graduate program gives you everything you need. (And don’t worry, there’s no training bond. No exit fees, no hidden catches). Here at mthree, we pair great graduates with brilliant global businesses. Our clients include tier one investment banks and other organizations across a range of industries, from insurance to healthcare to travel. As part of the mthree Alumni program, mthree has an exciting and exclusive partnership with Columbia Univ. School of Engineering. All mthree Alumni are eligible to receive two Executive Education certificates from Columbia Engineering as part of their Academy and industry placement experience at no cost. Further, all participating Alu
…What they are looking for
Skills & requirements
Department · Graduate - Technology
Hiring company
mthree
Explore this employer's active roles, salary signals and company profile on Jobiba.
Keep exploring
Similar active roles
Fresh roles matched to this title and market.
From $243.3K/yr
🔔 Get job alerts
New Site Reliability Engineer / Production Support Analyst jobs in Salt Lake City, Utah, United States, straight to your inbox.
No spam · Unsubscribe anytime