About AlphaSense: The world’s most sophisticated companies rely on AlphaSense to remove uncertainty from decision-making. With market intelligence and search built on proven AI, AlphaSense delivers insights that matter from content you can trust. Our universe of public and private content includes equity research, company filings, event transcripts, expert calls, news, trade journals, and clients’ own research content. The acquisition of Tegus by AlphaSense in 2024 advances our shared mission to empower professionals to make smarter decisions through AI-driven market intelligence. Together, AlphaSense and Tegus will accelerate growth, innovation, and content expansion, with complementary product and content capabilities that enable users to unearth even more comprehensive insights from thousands of content sets. Our platform is trusted by over 6,000 enterprise customers, including a majority of the S&P 500. Founded in 2011, AlphaSense is headquartered in New York City with more than 2,000 employees across the globe and offices in the U.S., U.K., Finland, India, Singapore, Canada, and Ireland. Come join us! About The Role: Our Site Reliability Engineering team is growing, and we are looking for a highly experienced Staff Site Reliability Engineer to help shape the future of reliability, scalability, and performance at AlphaSense. This is a hands-on, high-impact role where you will architect core reliability platforms, lead by example in incident response, and drive cultural adoption of SRE best practices across our global engineering organization. Our mission is to engineer our platform to the reliability standards of mission-critical systems, targeting 99.99% uptime, while continuously enhancing our systems and processes. This role is key to that mission and goes beyond traditional system maintenance; it’s about pioneering the platforms, practices, and culture that enable engineering to scale effectively. You will act as a force multiplier, mentoring fello
Jobiba hiring network
Site Director Jobs
1,315 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current site director jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an experienced Site Reliability Engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in an on-call rotation supporting highly available customer-facing systems. Lead incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with engineering teams to improve service availability, scalability, performance, and resilience. Continuously improve observability through metrics, logging, tracing, dashboards, and alerting. Eng
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team... GEDE (Global Edge and Domains Engineering) keeps GoDaddy's domains, edge, and aftermarket platforms running for millions of customers. Within GEDE, the Reliability Engineering team (Domains Production Engineering) is the group that gets the call when something breaks — and, more importantly, builds the systems that mean it breaks less often. We own the monitoring, compliance, and patching toolchain for the whole Domains infrastructure footprint, provide advanced incident support, and are actively modernizing how the org detects, diagnoses, and even auto-remediates issues with AI-assisted tooling! What you'll get to do... Build and evolve observability using Prometheus/Mimir, the Grafana LGTM stack, Elastic/OTEL, and Site24x7 — closing gaps across the org. Steer our cloud migration journey and bolster our efforts to keep the services reliable and performant. Own patching compliance and vulnerability remediation at scale across a mixed on-prem + AWS fleet, hitting hard SLA targets. Operate and extend our multi-tenant Kubernetes/ArgoCD platform, including the migration of core services. Contribute to our AI/automation initiatives: auto-generating runbooks from Prometheus alerts, ServiceNow change-risk scoring, and other tooling that reduces toil for the whole team Consult with partner dev teams on metrics, alert thresholds, and monitoring standards — this is a platform-enablement role, not just a ticket queue. Mentor other engineers on the team and help mature our operational practices. Your experience should include
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. At Okta, our motto is "Always On" and nowhere do we embrace that more than in Technical Operations. We strive to build the most reliable and performant systems on the planet through the skillful use of automation. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. You will work on: Leading the architecture, design and rollout of the internal developer platform which would span CI/CD, tooling, infrastructure as code (IAC) integrations as well as modernization efforts Enhance existing automation for fleet management and build new features as per the needs of the infrastructure platform Spearhead initiatives and projects that enhance engineer productivity by identifying and mitigating bottlenecks in the development flow Mentoring, managing, and leading a team of SWEs & SREs with a broad range of expertise and experience Triaging and troubleshooting complex production issues to ensure reliability and performance. Working closely with our stakeholders across the organization to ensure our new capabilities are aligned to our competing constraints of reliability, security, and delivery velocity Partnering directly with recruiting and people ops to hire and retain the best talent
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About The Team The Commerce Site Reliability Engineering team is responsible for the reliability, scalability, and day-to-day operation of the platforms that power GoDaddy's Commerce ecosystem. We build and operate shared infrastructure, support critical production systems, and partner closely with engineering teams to ensure services remain secure, resilient, and highly available. As a Senior Site Reliability Engineer, you'll join a team that values ownership, operational excellence, and continuous improvement. Engineers are empowered to identify problems, drive meaningful change, and influence how reliability is delivered across the broader Commerce organisation. From improving operational maturity and reducing toil to modernising delivery platforms and strengthening incident response practices, this team plays a key role in enabling engineering teams to move quickly and safely. You'll work closely with engineers across infrastructure, cloud, security, networking, and application teams while helping shape the future of reliability engineering at GoDaddy. This role offers significant opportunity to broaden your impact, develop technical leadership skills, and grow toward Staff and Principal engineering positions over time. What you'll get to do... Lead reliability and operational improvement initiatives across GoDaddy's Commerce platform, helping engineering teams build and operate services safely and at scale. Own critical production systems, drive incident response and post-incident improvements, and continuously raise the bar
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement b
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a key contributor within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. What You'll Be Doing Reliability & Operations Design, build, and operate large-scale cloud infrastructure and production services. Participate in a global on-call rotation supporting highly available customer-facing systems. Participate in incident response efforts and drive post-incident reviews focused on systemic improvements. Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Partner with en
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What you will be doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuous
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and continuously invest in platform engineering, observability, and operational excellence to enable our engineering teams to move quickly and safely. The Staff SRE, Classified Opportunity This role is ideal for an engineer who enjoys solving complex technical challenges at scale, building automation, and improving the reliability of production systems. You will serve as a technical leader within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build, and operate world-class cloud services. The ideal candidate exemplifies the philosophy of "if you have to do it more than once, automate it" and possesses a strong passion for continuous improvement, operational excellence, and software engineering. Security Clearance: Active U.S. TS/SCI clearance with Full Scope Poly Compliance Expertise: Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6) What You’ll Do Work with various teams to design and implement scalable, and reliable network solutions Maintain a highly available cloud infrastructure edge for the Okta identity platform C
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operati
Location Details: At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. About the Team Global Compute builds and operates the core cloud infrastructure that engineering teams rely on every day. We provision and manage AWS accounts across the company, operate the network backbone that connects them, and maintain the security guardrails that keep those environments safe, compliant, and scalable. We believe reliability is an engineering challenge, not an operations task. We automate repetitive work, build for scale before it becomes a problem, and invest heavily in observability to identify issues before they impact the business. What you'll get to do... Operate and scale AWS production infrastructure, owning the health of services that provision, secure, and manage accounts across GoDaddy AWS organisations. Design, build, and maintain cloud platform capabilities using Python, CloudFormation, AWS CDK, and automation-first practices. Drive cost optimisation initiatives that improve efficiency and deliver measurable business impact. Improve observability through monitoring, alerting, dashboards, and operational tooling. Participate in on-call rotations, lead incident response efforts, and drive long-term reliability improvements through blameless post-incident reviews. Support strategic AWS initiatives across networking, identity, governance, and multi-account architecture. Review code and designs, contribute documentation and operational runbooks, and mentor fellow engineers. Leverage AI-assisted tooling to improve engineering productivity, accelerate automation, and reduce operat
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. * This position requires the ability to access U.S. National Security information. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15 ) upon hire. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security &
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards. Incident Response: Participate in on-call rotations and be a primary responder for critical inci
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Staff DB SRE to build the runtime foundation for NVIDIA’s enterprise AI platforms — with a strong emphasis on database infrastructure at scale. This role blends large-scale database transformation with the building and development of GPU-accelerated platforms. You'll develop the software systems, automation frameworks, and high-performance database services that power NVIDIA’s AI workloads at scale. What you'll be doing: Design and operate highly available database clusters (MySQL, MSSQL, Oracle) with automated replication, failover, point-in-time recovery, and disaster-recovery strategies at enterprise scale. Drive database performance engineering — own query optimization, indexing strategies, connection pooling, lock-contention analysis, and storage-engine tuning for production systems handling millions of transactions. Build self-service database lifecycle automation — from one-click cluster provisioning and schema migrations to zero-downtime upgrades, blue-green deployments, and automated capacity scaling. Bridge relational and AI-native data infrastructure — extend traditional database exper
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a Senior Site Reliability Engineer – Storage, you will own the reliability, performance, and scalability of our global NAS, SAN, and Object Storage platforms that power critical internal and external services. You will combine deep storage expertise with strong automation and SRE practices to design, build, and operate highly available storage systems at scale. What You Will Be Doing: Lead design, deployment, and operations of production NAS, SAN, and Object Storage platforms, ensuring reliability, performance, and security. Capture requirements from partner teams, architect storage solutions, and drive end‑to‑end implementation for new and existing services. Develop, maintain, and improve automation for provisioning, configuration, monitoring, incident response, and lifecycle management of storage infrastructure. Participate in on‑call and incident response, lead troubleshooting of complex storage and performance issues, and drive root cause analysis and preventive actions. Define and track SLOs/SLIs and error budgets for storage services, using observability and analytics to continuously improve reliability and efficiency. Build and maintain runbooks, standard operating procedures, and comprehensive documentation for storage services and automation.<
Get new site director jobs by email
Daily job updates · Unsubscribe anytime