Jobiba hiring network

Site Lead Jobs

1,324 active opportunities · Updated for October 2026

Fresh results

4 shown

Explore current site lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
OpenAI
📍 San Francisco• Full-time
1mo ago

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About Role We are building a supply-chain organization capable of supporting our transition from laboratory-scale development to factory-scale operations. We are looking for an Inventory Manager to build and operate our inventory function from the ground up. This is a highly hands-on role. In the near term, you may start with an empty room and be responsible for determining what racks, shelving, bins, labels, scanners, workflows, and systems are needed to turn it into a functional stockroom. You will receive material, organize inventory, perform counts, move parts between buildings, resolve discrepancies, and establish the processes others will eventually follow. As we grow, you will have the opportunity to develop this foundation into a full-scale, multi-site factory inventory operation. You may hire and manage contractors or onsite inventory administrators, but this role will initially have a significant individual-contributor component and will remain accountable for day-to-day execution. This role owns inventory and internal logistics. It does not own purchasing, production planning, or inbound and outbound supplier logistics. This role is based in San Francisco, CA and requires in-person presence 5 days a week. In this role you will: Build inventory operations from the ground up across various OpenAI facilities. Design and set up stockrooms, receiving areas, and material-storage locations, including selecting racks, shelving, bins, carts, labeling equipment, scanners, and other infrastructure. Personally execute core inventory

awsrestai
View job →

About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int

pythonawsrest
View job →
O
1mo ago

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

pythonawsazure
View job →
O
OpenAI
📍 Washington• Full-time
1mo ago

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role Our technologies support some of the most important and impactful work in the world, including our strategic and high-impact customers in the public sector. As a Forward Deployed Security Engineer (FDSecE) you will be responsible for securing these novel applications of OpenAI’s technology. We’re looking for motivated, tenacious, and curious people who will work closely with engineering teams to ensure our infrastructure deployments are highly secure against our adversaries. As an FDSecE, you will embed directly throughout the lifecycle, working on-site and being hands-on to ensure the overall security of these deployments from design to production and through ongoing operations. This role is preferred to be based in Washington DC but may consider remote work. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel to and working from customer sites is required for this role. In this role, you will: Deeply embed with our most strategic public sector customers to implement and maintain robust security controls. Be a design and technical thought partner by leveraging security expertise on protective controls including access controls, authentication, encryption, network, and system security. Collaborate closely with teammates, cross-functional teams, customers, and service providers to achieve security and compliance goals. Ensure continuity of critical security and monitoring c

pythonawsazure
View job →
🔔

Get new site lead jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More site lead opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.