The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim
Jobiba hiring network
Senior Shift Engineer Jobs
7,292 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior shift engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur
We are fueled by a moral imperative to advance mankind, and it all begins with our people, our product, and our purpose. Passion isn’t something we turn on and off; it’s woven into everything we do. If you thrive in high-challenge environments, are inspired by exceptional teammates, and are driven to grow beyond what you thought possible, MX is where you belong. Come build the future with us. Join an award-winning company that isn’t just shaping the financial industry, but transforming it in ways that create meaningful, lasting impact for millions of people. At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime. We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand. We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric. This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought. Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL an
About the Role: We are looking for a Senior DevOps Engineer to join our DevOps team at K Health. You will own and evolve the infrastructure underpinning a healthcare AI platform serving patients and enterprise health system partners. This is a high-ownership role: you will architect and operate cloud environments across K Health and its enterprise partners, lead complex infrastructure migrations, drive disaster recovery programs, and help build the next generation of AI-powered operations tooling. You will also mentor junior engineers and collaborate closely with product and engineering teams across the company. This is a hybrid role based in New York City (4 days/week in office) and includes participation in a daytime on-call rotation. What you will do: Own the design, implementation, and evolution of our GKE-based Kubernetes infrastructure across K Health and enterprise partner environments. Build and maintain our Terraform modular infrastructure library, including reusable modules with automated testing, across GCP, Cloudflare, and AWS. Architect, build, and maintain GitLab CI/CD shared pipeline templates used by all engineering teams (build, test, security scanning, deployment). Own and maintain self-hosted infrastructure software running in-cluster, including GitLab, ArgoCD, Langfuse, DependencyTrack, NGINX Ingress, and others. Implement and support security and compliance controls across infrastructure and the software supply chain - secrets management, pipeline secret detection, container scanning, SOC2 and HIPAA. Drive disaster recovery readiness: design failover scenarios, author runbooks, and lead periodic DR tests. Lead development of AI-powered operations tooling and agentic infrastructure. Monitor, troubleshoot, and improve production system reliability; respond to incidents during on-call shifts. Mentor junior DevOps engineers and establish team-wide engineering standards. What we are looking for: 5+ years of experience in DevOps, platform engineering,
About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Responsibilities Lead high-impact platform projects — design and ship capabilities that move the needle on developer experience, reliability, or security, and set the bar for quality, testing, and safe deployment practices. Build the AI-augmented platform. Design tooling and workflows that help engineers get more out of AI-assisted development — think infra primitives that are easy to reason about, automated review, and policy-as-code that keeps the guardrails strong as AI shifts how code gets written. Own Infrastructure-as-Code for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — and make it consumable enough that an LLM can safely PR against it. Evolve our CI/CD backbone (Argo CD / Workflows / Rollouts, GitHub Actions) to make deploys faster, safer, and easier to reason about. Instrument and operate. Drive observability with Datadog and Amplitude, own dashboards and SLOs, and use the data to push reliability forward. Participate in on-call, lead incident response when needed, and turn postmortems into durable platform improvements. Reduce toil and tech debt with pragmatic remediation
We are a global team of innovators and pioneers dedicated to shaping the future of observability. At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first world by giving them unparalleled insight into their complex systems. As we continue to expand our global footprint, we're looking for passionate people to join our mission. If you're ready to help the world's best companies optimize their digital applications, we invite you to explore a career with us! Your Opportunity As a Senior Software Engineer within the Container Fabric (CF) organization, you will be a key driver in evolving New Relic’s global internal platform. We are looking for an operations-heavy engineer with 5–8 years of relevant experience who can leverage open-source and custom tooling to orchestrate and maintain large-scale Kubernetes environments. You will play a "Captain" role—leading critical deliverables and mentoring junior engineers while maintaining the reliability of our global fleet. What You'll Do Architectural Leadership: Drive the design and implementation of internal tools, specifically focusing on Kubernetes Operators and Controllers to automate resource management. Platform Orchestration: Lead complex, large-scale infrastructure shifts. Operational Excellence: Take ownership of incident response, author comprehensive retrospectives, and implement systemic hardening to prevent recurrence using advanced overcommit strategies. This Role Requires Experience: 5–8 years in a DevOps, Site Reliability, or Infrastructure Engineering role. Kubernetes Mastery: Deep internals knowledge of Kubernetes and hands-on experience writing custom operators. Tooling Proficiency: Strong experience building production-grade tools and services, specifically for infrastructure automation. Operations-Heavy Mindset: A proven track record of Day 1/Day 2 operations for a large-scale Kubernetes fleet, handling high-severity incidents, and improving SLA compliance through auto
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The Proactive Insights team is a newly formed team at the center of Mixpanel's AI-first analytics vision. With a greenfield charter, we're building the intelligent layer that transforms Mixpanel from a tool you query into a partner that works for you. We answer the question every data-driven team asks: "What changed, why, and what should I do about it?" We proactively keep users informed about what matters in their data, delivering the right insights and recommendations at the right time, to the right places, both inside and outside of Mixpanel. Some examples of what we are building: AI-powered root cause analysis : intelligent agents that diagnose why metrics changed, identify contributing factors, and recommend what to do next AI automations for KPI monitoring : always-on agents that track your key metrics and deliver insights to Slack, email, or directly in Mixpanel Intelligent alerts and anomaly detection : monitoring powered by models like TimesFM that surfaces meaningful shifts in your data About the Role As a Software Engineer on Proactive Insights, you'll build the features that transform how teams stay on top of their data. You'll work across the full stack, from React frontend experiences to Python backend services to integrating with our query engine, to ship workflows that surface insights users didn't know to ask for. You'll collaborate closely with Product and Design to shape what proactive analytics looks like, and with our AI Engine team to leverage shared infrastructure for orchestration, evaluation, and observability. This is a high-impact role on a small, fast-moving tea
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. * This position requires the ability to access U.S. National Security information. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15 ) upon hire. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security &
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Site Reliability Engineer (SRE) - Security and Data Systems Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts. Responsibilities Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems. Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response. Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards. Incident Response: Participate in on-call rotations and be a primary responder for critical inci
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Auth0 provides an unparalleled authentication experience for hundreds of millions of users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in Europe to ensure our production systems are not only operational but also resilient, scalable, and ready for exponential growth. This isn't just about keeping the lights on; it's about directly contributing to the platform's core resiliency and robustness. You'll be a hands-on builder, crafting solutions that make our system more reliable by design. What you’ll do: Design and build custom software in Go to enhance the platform's reliability, resiliency, and redundancy. Partner with engineering teams to embed reliability principles, improving the availability, performance, and observability of our services. Use your deep understanding of infrastructure and observability principles to identify opportunities for improvement within the product and implement solutions. Contribute to our follow-the-sun on-call rotation, providing rapid, effective response to critical incidents and using your expertise to troubleshoot, mitigate or accurately escalate production issues. Because our team is globally distributed, your on-call shifts will only occur during your standard local working hours. Develop and refine our SRE tooling and proc
The Developer Experience (DX) team at Amplitude builds and maintains the foundations that power how developers integrate, extend, and trust Amplitude across platforms. Our mission is to make Amplitude’s SDKs reliable, easy to adopt, and a joy to build on, so customers can confidently instrument their products and unlock insights at scale. We’re looking for a Staff Software Engineer, iOS to play a key technical leadership role on our DevEx team. In this role, you will lead the design and development of Amplitude’s core iOS SDKs, including Analytics and Session Replay , and serve as the iOS platform expert that other SDK teams, such as Statsig, Guides, and Surveys rely on. As a Staff Engineer, you’ll operate with a wide scope and high impact: setting technical direction for the iOS platform, driving cross-SDK architecture, improving performance and reliability, and raising the bar for developer experience across Amplitude’s mobile ecosystem. As a Staff Software Engineer, you will: Lead the technical direction, architecture, and long-term evolution of Amplitude’s iOS SDK platform. Own and drive development of core iOS SDKs, including Analytics and Session Replay, with a strong focus on performance, reliability, and ease of use. Act as the iOS platform expert and trusted partner for other SDK teams (Experiment, Guides, Surveys), enabling them to build on shared foundations safely and efficiently. Design and evolve shared infrastructure, APIs, and abstractions that scale across multiple iOS SDKs. Collaborate closely with Product, and Customer Support to ensure SDKs meet real customer needs. Lead cross-team technical discussions, reviews, and architectural decisions that span multiple SDKs. Improve developer experience through better APIs, documentation, tooling, testing strategies, and sample apps. Mentor senior and mid-level engineers, raising the overall quality and effectiveness of the team. You'll be a great addition to the team if you have: A strong focus on develop
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role The IQ Design team is dedicated to creating simple, delightful, and transformative user experiences. We’re passionate about developing innovative yet familiar interfaces that feel like natural extensions of the human experience. We value collaboration across disciplines—from software engineering to hardware design—and we’re excited to work with designers who see technology as a tool to help people tackle real-world problems. As a Senior Digital Product Designer, you’ll imagine, design, prototype, and ship compelling new AI experiences. The ideal candidate brings intense curiosity, a dedication to craft, a willingness to learn, and an openness to sharing and evolving ideas. You love experimenting, engaging deeply in the design process, and iterating continuously on concepts until they feel “just right” (and even then, you keep exploring ways to improve). At HP IQ, you’ll have the opportunity to develop features and products that people love using every day. What You Might Do Explore, design, and ship new experiences powered by AI Collaborate with fellow designers and engineers to refin
The Team + The Role Pendo’s Customer Engineering team is the technical backbone of the pre- and post-sales customer motion. The team brings together work across Customer Success, Technical Account Management, and Solutions Engineering into one full-lifecycle technical owner. Customer Engineering helps customers connect technical execution to business outcomes and realize the value of their investment in Pendo. As a Senior Customer Engineer for the ANZ region, you will own accounts across acquisition, implementation, adoption, expansion, and escalation. You will lead technical strategy for complex deployments, resolve ambiguous customer challenges, and partner closely with account teams to drive retention, expansion, and customer value. Your impact will extend beyond your own accounts through playbooks, coaching, repeatable fixes, and workflow improvements that raise the capability of the broader Customer Engineering team. This role is based in our Sydney office. What this looks like day-to-day New customer acquisition and selling: Partner with Account Directors on pre-sales motions for new customer acquisition, product expansion, and growth across lines of business. You identify customer pain, craft and deliver tailored demos, and lead technical evaluations through requirements gathering, success criteria, installation guidance, and hands-on workshops. Quarterly business outcomes: Drive measurable improvement in customer health, retention, expansion, and technical resolution across strategic accounts. You contribute at least one reusable framework, playbook, or process that other Customer Engineers use and become a trusted advisor to Account Directors on account health and technical strategy. Adoption, expansion, and growth: Identify opportunities to expand adoption across underutilized features, untapped products, and new lines of business. When adoption stalls, you diagnose root causes, design a path forward, and connect technical work to measurable business outco
Technical Services operates with a global footprint, maintaining a presence in key cities such as New York City, Toronto, Austin, Palo Alto, Vancouver, Sydney, Gurgaon (India), Tel Aviv, Dublin, and Buenos Aires. Our commitment to exceptional customer satisfaction is underpinned by a 24x7x365 'follow-the-sun' support strategy, executed by dedicated regional teams across the Americas, EMEA, and APAC. We are seeking a visionary leader to join the MongoDB Technical Services organization to oversee regional teams of MongoDB Premium Services Support Engineers. These specialists possess deep expertise in resolving challenges related to core database functionality and troubleshooting complex cloud environments to facilitate high-scale customer success. For this hybrid role, we are specifically interested in meeting candidates located in Palo Alto, CA or San Francisco, CA. Core Responsibilities of the Team Diagnosing and resolving performance issues: Identifying, troubleshooting, and fixing performance-related bottlenecks Advising on design and architecture: Providing technical guidance on architectural design, including replication and global data distribution, to ensure low latency, high availability, and compliance with data sovereignty requirements Serving as a customer advocate: Advising on upcoming roadmap features and articulating strong business cases to drive the swift resolution of critical issues Engaging in global collaboration: Partnering with regional peer teams and sharing account context to anticipate client needs before an event, ensuring a seamless and elevated level of global service Fostering a culture of continuous learning: Growing technical skills through training participation, maintaining industry awareness, and developing specialize
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor
Get new senior shift engineer jobs by email
Daily job updates · Unsubscribe anytime