Jobiba hiring network

Site Lead Jobs

1,324 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current site lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. EXAMPLE INITIATIVES You'll work on projects like these as part of the SRE team: Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services. Building AI-assisted tooling for incident triage and response. RESPONSIBILITIES Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking. Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code. Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution. Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations. Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, con

kubernetesgitmachine learning
View job →
S
Supabase
📍 Remote• Full-time
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform. You'll be embedded within Service Operations, and your primary job is to make every engineering team more reliable — not by owning their infrastructure, but by establishing the practices, frameworks, and feedback loops that let them own reliability themselves. You'll work across the org: sometimes setting the standard, sometimes pair-programming a fix, sometimes helping a team define their error budget, sometimes telling them it's exhausted. This role is ideal for someone who has a strong vision for how SRE should work and thrives in async, fast-paced environments where influence matters more than authority. What You'll Own Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience, and build the error budget policies that turn them into engineering decisions Own and evolve the Operational Readiness Review (ORR) process — conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes Act as the reliability expert teams pull in for architecture reviews, failure mode analysis, dependency mapping, and resilience design Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it

awskubernetesai
View job →
G
Godaddy
📍 United States• Full-time• From $128K/yr
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor

pythonkuberneteslinux
View job →
M
Mongodb
📍 Toronto• Full-time• From C$144K/yr
1mo ago

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based

pythonmongodbaws
View job →
M
Mongodb
📍 United States• Full-time• From $127K/yr
1mo ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

mongodbawsazure
View job →

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture. This relatively new team is building performant, multi-tenant distributed storage services that both enhance today’s Atlas storage stack and enable more customer workloads to run more efficiently. You will partner with the teams building these storage services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins Atlas. You’ll join a small, senior team of SREs as founding members of this organization, playing a crucial role in executing on a multi-year roadmap for MongoDB’s cloud storage architecture. This role can be based out of our Boston, New York City, Raleigh, Miami, Pittsburgh or remotely in the United States while physically based in an Eastern or Central time zone location. The ideal candidate should Have 6+ years of experience working on software development and operating distributed systems Proficiency in Python, Go, or a similar language Have operated or supported stateful storage or database systems at scale, and are comfortable with durability, consistency, and recovery trade-offs. Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Responsibilities Work on our multi-tenant distributed storage systems, balancing long-term strategic infrastructure g

pythonmongodbaws
View job →
M
1mo ago

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur

mongodbawsazure
View job →
M
1mo ago

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Deployments team designs and maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering teams. This infrastructure is primarily composed of Argo Workflows and ArgoCD. The team also provides tooling that enables clear system ownership and facilitates self-service onboarding for development teams. We are looking to speak to candidates who can work East Coast hours. The ideal candidate should Have 6+ years of experience in software development and operating distributed systems Proficiency in Python, Go, or a similar language Proven experience building and operating large-scale continuous integration and continuous deployment (CI/CD) pipelines Possess a customer-focused mindset Value efficiency in processes and operations Prefer automation over manual process (“allergic to ops work”). We are a small team of software engineers with a strong bias towards software solutions to avoid toil Experience using and extending containerization technologies, particularly Kubernetes, to enhance application agility, optimize resource utilization, and accelerate time-to-market Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure Understanding of Linux operating system internals and networking concepts (e.g., TCP/IP, DNS, TLS, routing) Expectations Contribute to developing a world-class continuous deployment experience, enabling the rapid and reliable shipment of MongoDB products This includes, but is not limited to, contributing to open-source projects, or engineering software-based

pythonmongodbaws
View job →

The Team This role can sit in our NYC HQ on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas platform. As a senior SRE, you will be expected to be able to design & build complex systems, operate with autonomy and act as owner for everything you do. The SRE Atlas team works alongside the various Atlas software engineering teams to provide expertise about running systems at scale, build new tooling and automation and perform essential maintenance of the Atlas fleet. This is an SRE team, which means you can expect a highly hands-on approach, tackling the technical challenges of implementing large scale solutions that have the ability to impact our customer’s most crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This role requires engineers to have a customer-first mindset to ensure that everything we do results in a stronger product and a better experience for all Atlas customers. The ideal candidate should Have 5+ years of experience running critical systems at scale Value efficiency in processes and operations, and display a preference for automation over manual processes (“allergic to ops work”) Be familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment A strong understanding of how to run a large scale Linux environment, including low level fundamentals Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python) Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc) Special Requirements: Be a US Citizen Expectations Participate in the development of a reliable and resilient multi-cloud platform that hosts business critic

pythonmongodbaws
View job →
🔔

Get new site lead jobs by email

Daily job updates · Unsubscribe anytime

Explore verified demand

More site lead opportunities

Browse all jobs →

Companies hiring

Employers are derived from current jobs in this exact search market.