Jobiba hiring network

Sre Operations Engineer Jobs

197 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current sre operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

D
Datadog
📍 California• Full-time• Remote
1mo ago

We are a team of engineers that translate our real-world experience to help our user communities solve problems. With a focus on service management, helping teams respond to incidents, run on-call, and automate their operations, you will work with practitioners and leaders across the industry and broaden your impact to the SRE, Engineer, DevOps, and Operations community at large. This is a unique opportunity to use both your engineering and creative storytelling skills to shape the landscape in cloud observability, incident response and service management. What You'll Do: Act as a subject matter expert for service management (incident response, on-call, IDP, Work Management, Workflow Automation, Agent Builder, and operational automation) for Datadog's advocacy and engineering teams Create content in one or more mediums to build Datadog's reputation as a leader in DevOps, Monitoring, Observability and Security e.g. building demos, public speaking, blogging, documentation, webinars, open source, research reports and more Partner with product engineering teams to build compelling demos, and coach internal engineering teams on effective communication and presentation Interface with open source communities to drive key messaging in the market and develop new integrations for Datadog Contribute to the product through feedback (bugs or product enhancements suggestions), documentation, or code Who You Are: Approximately 5+ years of experience as a Platform Engineer, Site Reliability Engineer, DevOps Engineer or Software Developer with hands-on experience as an on-call/incident responder and running production systems in complex IT environments You have a strong understanding of core service-management practices (incident response, on-call, post incident reviews, and SLOs), using tools like Datadog, PagerDuty, Opsgenie, incident.io, Rootly, Jira Cloud Platform, Cortex, or similar and know how to navigate operational challenges of different s

REMOTEpythonnode.jsai
View job →
P
Pagerduty
📍 Remote (USA)• Full-time• Remote
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. PagerDuty is seeking a Principal Solutions Consultant to join our talented, customer-focused team! You will be a key strategic leader and the "technical face" of PagerDuty across the US region. This is a high-impact role that sits at the intersection of Sales, Product, and Engineering. You won’t just be selling a platform; you will be partnering with CTOs, CIOs, and VPs of Engineering at the US’s most influential enterprises to redefine how they manage digital operations, resilience, and automation. You will act as a bridge between our customers’ long-term strategic needs and PagerDuty’s product roadmap. You will spend your time evangelizing the PagerDuty Operations Cloud, mentoring our high-performing technical sales teams, and acting as a trusted advisor to the C-suite on topics ranging from AIOps and SRE maturity to digital transformation and cloud migration. Key Responsibilities Act as a peer and advisor to customer CTOs and CIOs, helping them navigate complex digital transformations and operationalizing the PagerDuty Operations Cloud within their organizations. Represent PagerDuty as a thought lea

REMOTEawsgitrest
View job →
D
Datadog
📍 New York• Full-time• From $156K/yr
21 days ago

As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to

gitairust
View job →
D
1mo ago

We are Datadog's in-house product experts. The Technical Solutions team enables Datadog’s worldwide growth by educating potential clients and ensuring that existing customers are happy and successful. As a Technical Account Manager 2 (TAM 2), you’ll serve as a trusted advisor to our strategic customers, accelerating their adoption of the Datadog platform and enabling long-term success. TAM 2s bring deep technical expertise, refined customer skills, and consultative insight into how monitoring, observability, and DevOps practices translate to business value. At Datadog, we place value in our office culture—the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Act as a technical advisor to 3 primary enterprise accounts, ensuring successful product adoption and effective usage of the Datadog platform. Lead enablement and adoption sessions across core product areas tailored to your customer’s architecture and business needs. Analyze customers’ IT Operations environments and workflows to recommend configuration, product usage, and performance improvements. Deliver technical business reviews, health checks, and account maturity assessments, contributing to customer QBRs with impactful recommendations. Escalate product issues appropriately and advocate for your customers’ needs with Datadog’s Product and Engineering teams. Create executive-level summaries and insights that tie platform usage to business outcomes. Participate in internal TAM strategy sessions and contribute to team learning through feature presentations, case studies, or best-practice sharing Who You Are: You have 2+ years of experience in a technical customer-facing role (TAM, Solutions Architect, SRE, DevOps Engineer, or similar) within the cloud or observability space. You’re confident with at least two public cloud platforms (e.g. AWS, Azure, GCP)

pythonawsazure
View job →
N
9 days ago

As a Software Solution Architect, NVIS at NVIDIA, you will lead the transformation of AI infrastructure. Our NVIS team focuses on developing the next generation of NVIS Central, an agentic software platform with tools, services, and AI agents that automate, simplify, and speed up the work of our delivery organization. This role offers an outstanding chance to create and build LLM-powered agents that improve execution visibility, cut down manual tasks, and standardize workflows. These efforts allow NVIS to grow quickly and with high quality. Join us to bring up, validate, optimize, and upgrade large-scale AI Factory infrastructure for some of the world’s most advanced accelerated computing environments! What you'll be doing: Compose, build, and productionize agentic AI solutions, tools, and applications for the NVIS delivery organization. Develop LLM-based agents, skills, tool-calling workflows, orchestration logic, backend services, APIs, data pipelines, and automation features as part of NVIS Central. Translate field, delivery, operations, and product needs into clear technical builds, agent workflows, and working software. Develop agents that can reason across project data, knowledge bases, operational systems, logs, reports, and delivery workflows. Build workflows that help NVIS teams identify risks, summarize project status, automate repetitive tasks, improve readiness visibility, and simplify handoffs. Work with timely engineering, retrieval-augmented generation, context management, agent memory, function calling, evaluations, and guardrails to build reliable AI systems. Integrate LLMs and agents with internal systems, project data sources, knowledge repositories, reporting tools, and operational workflows. Collaborate closely with software developers, architects, product managers, DevOps/SRE, and NVIS field teams to

pythonsqldocker
View job →
G
Godaddy
📍 Bulgaria• Full-time
1mo ago

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. Remote: This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Our Global Sustaining Engineering team sits at the intersection of software engineering and infrastructure, ensuring the services our customers depend on are fast, resilient, and always available. As a Senior Site Reliability Engineer, you'll take direct ownership of production services — from initial design through day-to-day operation — while partnering with product, engineering, and security teams to build and maintain business-critical systems. In this role, you will deepen your technical expertise and grow your leadership presence by mentoring the next generation of SREs. You will also gain hands-on experience with intelligent tooling in real-world workflows. What you'll get to do... Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana Lead incident response end-to-end — performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure Develop and extend Infrastructure as Code coverage and build internal tooling that eliminates manual, repetitive operational work Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation — validating all outputs before use Your experien

pythondockerkubernetes
View job →
O
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Infrastructure Platform and Shared Services Team Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling. As the Sr. Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, Observability, automation platform & tooling. What you’ll be doing Lead the Infra platform and shared services org and various initiatives across SRE & Infrastructure organization. Build a world-class observability platform and monitoring capabilities enabled with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and intuitive self-service capabilities. Own the design and operation of scalable, self-service Cloud infrastructure platforms (e.g. Observability Platform, SRE Productivity, deployments, and Edge Infrastructure) Lead, mentor, and grow a high-performing team of engineers and managers across SRE and infrastructure shared services domains. Perform engineering design evaluations and ensure the completion of projects within resource,

awsci/cdrest
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $243.9K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Core Infrastructure team within Coinbase's Platform product group builds the foundational systems that keep Coinbase online, secure, and scalable, owning the compute and networking platforms that power every product and service across the company. As the Group Product Manager for Core Infrastructure & Reliability, you'll own the product vision and multi-year strategy for Coinbase's cloud infrastructure, driving the design, operation, and scaling of the systems that underpin hundreds of billions of dollars in annual transaction volume. You'll partner deeply with Engineering, SRE, Security, and Finance to ensure Coinbase's infrastructure is reliable, cost-efficient, and resilient across multiple cloud environments and regions. What you’ll do: Own the product strategy and roadmap for Core Infrastructure, spanning compute, networking, multi-region and multi-cloud architecture, and platform reliability. Strengthen infrastructure reliability and resilience programs, defining platform-level SLOs, capacity planning, failover capabilities, and incident reduction targets to meet the uptime demands of a global financial platform. Lead evaluation and adoption of cloud infrastructure technologies (Kubernetes, service mesh, distributed storage, observability, infrastructure-as-code), making build-vs-buy decisions that balance cost, speed, and long-term scalability. A

REMOTEawskubernetesai
View job →
S
Synthesia
📍 United States• Full-time
1mo ago

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp

pythonmongodbaws
View job →
R
14 days ago

Reolink , a leader in intelligent visual technology for homes and businesses, was founded in 2009 by a group of engineers with a strong commitment to and passion for smarter security solutions. Our products are now trusted by millions of users across more than 110 countries and regions worldwide. Building on this trust, we continue expanding our presence and bringing our innovations to more markets around the globe. Reolink remains committed to delivering advanced, reliable, and user‑centric solutions that empower people to protect what matters most. 5 Work Days Per Week Office at Tai Seng Exchange Tower B Near Tai Seng MRT, Singapore Insurance Coverage Entitled to Yearly Bonus & Performance Bonus Responsibilities (Site Reliability Engineer - Senior / Lead ) High Availability and Stability Maintenance of Application Systems: Includes daily monitoring, alert response, emergency handling, on-call duties, regular system health checks, and performance optimization. Compliance and Secure Access Construction for Application Systems: Ensure operational design, processes, and data management comply with relevant privacy and data protection laws. Ensure compliance with full auditing and regulatory checks and provide auditing materials as required. Change and Release Management: Best practices for application system changes, including change control, version management, and rollback strategies, while ensuring operational duties during release windows. Automation and Infrastructure Optimization: Drive operational automation by designing and implementing automated tools and processes, ensuring resource allocation is optimized and supporting business scalability. Other Operational Practices and Work Arrangements: Providefeedback and suggestions for business architecture design and continuously produce operational technical documentation. Qualifications Bachelor's Degree or above; a degree in computer science or a related field is preferred. Experiences as Senior SRE or

awsazuredocker
View job →
A
Affirm
📍 Poland• Full-time• Remote• $384K – $576K/yr
15 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →
A
Affirm
📍 Spain• Full-time• Remote• From €1M/yr
15 days ago

At Affirm, we exist for the moments that matter—giving people a clear, predictable way to pay over time, with no hidden fees, no surprises, and no tradeoffs on what matters most. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to “Operate What They Own” with excellence to protect their customers’ experience. SRE accomplishes this through defining frameworks and best practices for operating applications, building tooling, and providing training and consulting. Some of the many SRE responsibilities are: Providing data and visibility to teams and leadership on application performance Guiding the development of SLOs Driving the Incident Management and Analysis process Steering the implementation of Change Management and Deployment practices Engaging in service and architectural conversations Recommending observability and alerting configurations The SRE team benefits from experience across many domains including: infrastructure, platform, and distributed systems capacity management, load and chaos testing automation, observability, and configuration management development and product experience The SRE team is seeking motivated software and systems engineers with the experience to build, iterate on, and expand incident lifecycle, reliability, and resilience practices throughout Affirms Engineering organization and beyond. What You'll Do: You will be responsible for owning and delivering quarterly goals for your team, leading engineers on your team through ambiguity to solve open-ended problems, and ensuring that everyone is supported throughout delivery. You will support your peers and stakeholders in the product development lifecycle by collaborating with infrastructure, product management, developer experience & analytics by participating in ideation, articulating technical constraints, and partnering on decisions that properly consider risks and trade-offs. You will proactively identify technical solutions

REMOTEpythonsqlmysql
View job →

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is seeking a strategic and results-driven engineering leader who is passionate about cloud agnostic infrastructure, operational excellence, and enabling engineering teams to operate autonomously and build with confidence. As Head of Infrastructure, you'll lead a talented and geographically distributed team of engineers across the SF Bay Area, India, and Europe, fostering a culture of collaboration, ownership, and continuous improvement. You'll own the infrastructure that underpins one of the world's most widely used API platforms, an environment handling ~80,000 requests per second at the front door, and be responsible for its reliability, scalability, and evolution. In addition to infrastructure, you'll own the Site Reliability Engineering (SRE) function at Postman, setting the standards and practices that keep the platform reliable at scale. You'll work closely with engineering managers, product managers, and platform teams to drive the technical roadmap for our cloud agnostic infrastructure and reliability practices, ensuring we can support a large and rapidly growing engineering organization. If you're p

awsazurekubernetes
View job →
R
Remote
📍 EMEA• Remote
9 days ago

About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage international teams. With our core values at heart and future focused work culture, our team works tirelessly on ambitious problems, asynchronously, around the world. You can find Remoters working from 6 different continents (Antarctica left to go!) and all of our positions are fully remote. With Innovation as one of the core values, we have built Automation and AI capabilities into the requirements for every role. We encourage every member of the Remote team to bring their talents, experiences and culture to the table to help us build the best-in-class HR platform. If you are energetic, curious, motivated and ambitious, be part of our world. Apply now and define the future of work! What this job can offer you Remote's SRE team exists so that our engineers can move quickly and our customers get a product that stays up. The team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, our observability stack and the reliability practices that sits on top of all of it. We are looking for a Team Leader to run that team. This is a 60% IC, 40% leadership role. You will own the career development of your reports, steer the teams focus using judgment against the company goals, and you will be the spokesperson for the team across engineering. You will also stay close enough to the technical work to set direction with credibility and to know when something is going wrong before it is escalated to you. Reliability practice at Remote is maturing rather than mature. Our SLO framework is live on its first few teams and needs to reach the rest, there is real work to do on how we balance operational load against project delivery. If you want a team where the foundations are in place and the interesting problems are still open, this is that team. What you bring People leadership Yo

REMOTEpythonjavanode.js
View job →
P
Pinterest
📍 CA, United States• Full-time• Remote
1mo ago

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Production Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams' capability to design, build and operate robust systems at scale. Pinterest's applications and infrastructure handle billions of monthly page views and petabytes of data as Pinterest continues to grow and scale. As a Senior Production Engineer on Solutions Engineering, you will design and build AI agents, platforms, tools, frameworks and methodologies to assure the reliability of our large-scale distributed systems serving hundreds of millions of monthly active users, handling hundreds of thousands of requests per second, and managing tens of petabytes of data. You'll lead infrastructure modernization initiatives, build intelligent automation that eliminates operational toil and amplifies engineer

REMOTEpythonsqlmysql
View job →
🔔

Get new sre operations engineer jobs by email

Daily job updates · Unsubscribe anytime