Jobs in United States

Site Lead in United States

256 active opportunities · Updated October 2026

Explore current site lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 Washington, District of Columbia, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) which benefits all of humanity. This long-term undertaking brings the world’s best scientists, engineers, and business professionals into one lab together to accomplish this. In pursuit of this mission, our Go To Market (GTM) team is responsible for helping customers learn how to leverage and deploy our highly capable AI products across their organization. The team comprises Sales, Solutions, Support, Marketing, and Partnership professionals who collaborate to create valuable solutions that will help bring AI to as many users as possible. About the Role Our Government Sales team has a unique mission to help the U.S. Intelligence Community understand the transformative impact that highly capable AI models can bring to its most critical missions. This role combines technical understanding, strategic vision, relationship management, and value-driven sales strategy tailored specifically to Intelligence Community customers. You’ll drive key opportunities throughout the full sales cycle, from pipeline generation through deployment and expansion. You’ll collaborate closely with researchers, engineers, solution strategists, and cross-functional partners to help Intelligence Community customers advance their missions through AI. This role is based in Washington, D.C. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In addition, this role requires frequent on-site engagement at customer and partner facilities across the Washington, D.C./Northern Virginia corridor, including classified environments. In this role, you will: Manage a focused set of key U.S. Intelligence Community accounts, developing and executing comprehensive account strategies. Lead Intelligence Community customers through their AI adoption journey, from initial consideration through successful deployment and expansion. Build trusted relationships with senior

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is helping build the infrastructure that powers the next generation of artificial intelligence. Through Stargate, we are developing and operating large-scale AI compute campuses that require world-class execution across data center design, construction, commissioning, and operations. The Infrastructure Operations team is responsible for bringing AI infrastructure online and ensuring it operates reliably at scale. We partner closely with hardware, network, deployment, construction, and operations teams to deliver mission-critical environments capable of supporting frontier AI workloads. As our footprint expands, operational excellence becomes increasingly important to ensuring safe, reliable, and efficient campus operations. About the Role We are seeking a Facilities Operations Manager to support the commissioning, operational readiness, and long-term operation of next-generation AI data center campuses. This role sits at the intersection of construction, commissioning, hardware deployment, and facilities operations. You will be responsible for ensuring mission-critical infrastructure is prepared to support hardware deployment, transitioned successfully into production operations, and maintained to the highest standards of reliability and availability. You will lead day-to-day operational execution across electrical, mechanical, controls, and supporting infrastructure systems while partnering closely with commissioning teams, site operators, vendors, and engineering organizations. This role requires a strong blend of technical depth, operational leadership, and cross-functional execution. Key Responsibilities Lead day-to-day operations of mission-critical facility infrastructure across AI compute campuses. Own operational readiness activities supporting new campus deployments and infrastructure expansion. Partner with commissioning teams to transition facilities from construction and startup into steady-state operations. Develop, implement, and

AWSRestAIRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

KubernetesGitMachine LearningAI
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $128K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor

PythonKubernetesLinuxAI
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $127K/yr

Quick readStrong listing-quality and freshness signals

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

MongoDBAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o

AWSRestAIRust
Z
📍 Bellevue, Washington, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h

PythonKubernetesLinuxAI
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About Plaid & Our Mission At Plaid, we are building the digital financial ecosystem of the future, giving millions of people the power and access to manage their financial lives seamlessly. Behind every great product is a team of people who need reliable, friction-free technology to do their best work. Dynamics & Your Impact As part of the TechOps team, you will design, build, and support a first-class computing environment that keeps our business secure and our employees productive. In this role, you will be the backbone of our in-office technical experience. By ensuring our conference rooms, networking systems, and daily workplace tools run smoothly without interruption, you enable every employee to focus on expanding financial access for everyone. What You’ll Do (Responsibilities & Impact) Enable Seamless Collaboration: Ensure audio-visual (AV) systems in conference rooms and shared office spaces function flawlessly so team members can connect seamlessly across global offices. Resolve In-Office Hardware & Tech Needs: Directly troubleshoot and resolve issues with printers, monitors, keyboards, and other office desk peripherals, minimizing downtime for your collea

P
📍 Seattle, Washington, United States· Full-time
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About Plaid & Our Mission At Plaid, we are building the digital financial ecosystem of the future, giving millions of people the power and access to manage their financial lives seamlessly. Behind every great product is a team of people who need reliable, friction-free technology to do their best work. Dynamics & Your Impact As part of the TechOps team, you will design, build, and support a first-class computing environment that keeps our business secure and our employees productive. In this role, you will be the backbone of our in-office technical experience. By ensuring our conference rooms, networking systems, and daily workplace tools run smoothly without interruption, you enable every employee to focus on expanding financial access for everyone. What You’ll Do (Responsibilities & Impact) Enable Seamless Collaboration: Ensure audio-visual (AV) systems in conference rooms and shared office spaces function flawlessly so team members can connect seamlessly across global offices. Resolve In-Office Hardware & Tech Needs: Directly troubleshoot and resolve issues with printers, monitors, keyboards, and other office desk peripherals, minimizing downtime for your collea

R
📍 San Fransisco, California, United States· Full-time
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role We’re looking for an IT Site Specialist to join our IT Team at Ramp based in our San Francisco office! This is a contract on-site role that blends hands-on, in-person site support with ownership of our core SaaS application stack, and we weigh both sides equally. On the site side, you’ll be the primary IT presence in SF — owning deskside support, onboarding/offboarding, endpoint and AV support, and office inventory. On the application side, you’ll administer the tools the whole company depends on — Okta, Google Workspace, Slack, JAMF, and a growing set of SaaS platforms — owning provisioning, access, portfolio, renewals, integrations, and automation. We’re looking for someone who takes pride in running a smooth on-the-ground operation while also thinking like a systems owner who scales IT through automation and sound identity practices. What You’ll Do Provide in-person support at our San Francisco office 5 days/week , acting as the site’s primary IT point of contact. Perform IT Support Specialist (L4) duties, encompassing all responsi

PythonRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service discovery, secrets management and related software layers. We’re looking for a skilled Senior Site Reliability Engineer with strong programming skills to help us build Roblox's private cloud, productionize our growing Kubernetes-based infrastructure, and institute reliability best practices across the Roblox Compute team. You will: Design and Develop systems & libraries that promote fault-tolerance and resilience, automate much of the management and lifecycle of our clusters, and ensure systems are observable. Promote and Institute reliability best practices across the Infra Compute group, drive common reliability initiatives. Provides collaborative technical reviews and operational guidance to strengthen system reliability. Build, Automate and Standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem. Create Tooling that provides production guardrails, by evaluating release candidate capacity with load testing tooling before de

JavaAWSKubernetesGit
S
📍 Humacao 1200 Dr John A Smith Street, Catalio Ward L-140, United States
✓ High-confidence listingCompany trend +364.7%
Quick readStrong listing-quality and freshness signals

Work Flexibility: Onsite Como Plastic Processing Operator III , serás responsable de operar, configurar y monitorear maquinaria de producción, asegurando el cumplimiento de los estándares de calidad, seguridad y productividad establecidos. ¿Qué harás? Opera equipos de procesamiento de plásticos, incluyendo máquinas automáticas y semiautomáticas, para cumplir con los objetivos de producción y calidad. Configura, inicia y monitorea equipos de manufactura siguiendo hojas de proceso, procedimientos establecidos e instrucciones operativas. Realiza cambios de herramientas, moldes y ajustes de configuración para garantizar las dimensiones y especificaciones requeridas del producto. Controla parámetros del proceso, como temperatura, presión y velocidad, y ejecuta ajustes rutinarios para mantener una producción consistente. Inspecciona piezas mediante controles visuales e instrumentos de medición para verificar el cumplimiento de estándares dimensionales y de calidad. Registra con precisión datos de producción, desperdicio, ajustes de proceso y actividades básicas de mantenimiento. Ejecuta mantenimiento preventivo básico y reparaciones menores, escalando incidencias complejas al equipo de mantenimiento correspondiente. Manipula y empaqueta componentes terminados, manteniendo el cumplimiento de los procedimientos de seguridad, calidad y uso de equipos de protección personal. ¿Qué necesitas? Requerido: Diploma de cuarto año completo o equivalente. Experiencia mínima de 2 años en operación de maquinaria de manufactura

S
📍 Humacao 1200 Dr John A Smith Street, Catalio Ward L-140, United States
✓ High-confidence listingCompany trend +364.7%
Quick readStrong listing-quality and freshness signals

Work Flexibility: Onsite Como Plastic Processing Operator III , serás responsable de operar, configurar y monitorear maquinaria de producción, asegurando el cumplimiento de los estándares de calidad, seguridad y productividad establecidos. ¿Qué harás? Opera equipos de procesamiento de plásticos, incluyendo máquinas automáticas y semiautomáticas, para cumplir con los objetivos de producción y calidad. Configura, inicia y monitorea equipos de manufactura siguiendo hojas de proceso, procedimientos establecidos e instrucciones operativas. Realiza cambios de herramientas, moldes y ajustes de configuración para garantizar las dimensiones y especificaciones requeridas del producto. Controla parámetros del proceso, como temperatura, presión y velocidad, y ejecuta ajustes rutinarios para mantener una producción consistente. Inspecciona piezas mediante controles visuales e instrumentos de medición para verificar el cumplimiento de estándares dimensionales y de calidad. Registra con precisión datos de producción, desperdicio, ajustes de proceso y actividades básicas de mantenimiento. Ejecuta mantenimiento preventivo básico y reparaciones menores, escalando incidencias complejas al equipo de mantenimiento correspondiente. Manipula y empaqueta componentes terminados, manteniendo el cumplimiento de los procedimientos de seguridad, calidad y uso de equipos de protección personal. ¿Qué necesitas? Requerido: Diploma de cuarto año completo o equivalente. Experiencia mínima de 2 años en operación de maquinaria de manufactura o procesamiento industrial. Conocimiento básico de interpretación de planos, diagramas o esquemas técnicos. Capacidad para utilizar herrami

S
📍 Kalamazoo, Michigan, United States
✓ High-confidence listingCompany trend +364.7%

From $22/hr

Quick readStrong listing-quality and freshness signals

Work Flexibility: Onsite 3rd shift: Sunday - Thursday 10:00 pm - 6:30 am What You Will Do: Under general supervision, operate machinery and inspect machined components using precision measuring equipment while keeping accurate production records and maintenance logs Adhere to site specific quality systems and processes Identify and accurately record scrap, maintenance requests, and production documents Operate simple manufacturing equipment, demonstrate machining/mechanical aptitude, and learn new responsibilities and tasks as needed Train others on operational and/or documentation procedures as needed Identify and appropriately report safety concerns, production issues, and documentation errors. Assist execution of continuous improvement projects. Effectively collaborate with peers, functional departments, and visitors of Stryker What You Need: Preferred High School or GED 2+ Years of Manufacturing Experience Blueprint reading, measuring tools – calipers, micrometers, gauges $21.90 + $1.50 Shift Premium per hour plus bonus eligible + benefits. Travel Percentage: None Stryker Corporation is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, ethnicity, color, religion, sex, gender identity, sexual orientation, national origin, disability, or protected veteran status. Stryker is an EO employer – M/F/Veteran/Disability. Stryker Corporation will not discharge or in any other manner discriminate

H
📍 Arizona, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$59.4K – $89.6K/yr

Quick readStrong listing-quality and freshness signals

Logistics Analyst Description - Location: Fulltime onsite Litchfield Park/Phoenix Arizona Role Summary The Logistics Analyst is the on-site expert who orchestrates all shipping activities- picking, packing, staging, loading, and system transactions within HP’s distribution operations. Blending shop-floor coordination with advanced analytics, the role ensures that every order departs safely, on time, and in full. Working side-by-side with 3PL partners, Transportation, Customer Support, and IT teams, this individual contributor drives day-to-day execution while uncovering process improvements that elevate customer experience, cost efficiency, and compliance. This position will work onsite at our central DC in Litchfield Park/Phoenix Arizona. Key Responsibilities Oversee daily inbound or outbound operations, ensuring accurate and timely order fulfillment, packing, and shipping in line with HP’s quality and service standards. Lead and support process improvement initiatives focused on optimizing outbound flow, reducing order cycle times, minimizing shipping errors, and enhancing on-time delivery performance. Collaborate with Supply Chain, Customer Service, Warehouse, IT, and Transportation teams to coordinate outbound shipments, resolve exceptions, and ensure alignment with customer requirements. Track real-time metrics such as pick accuracy, order fill rate, on-time ship, trailer utilization, and dock-to-door cycle time; initiate immediate countermeasures when performance drifts. Serve as the primary point of contact for order expedites, configuration changes, embargo checks, and any customer-facing shipment issues. Develop and maintain SOPs for outbound operations, ensuring compliance with company policies and regulatory requirements.

SapExcelPower BiTableau
🔔

Get new site lead jobs in United States by email

Daily job updates · Unsubscribe anytime