Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance. You will: Design and Implement Observability Solutions : Develop comprehensive monitoring and alerting systems using modern observability tools. Create dashboards and metrics that provide real-time visibility into system health and performance. Implement logging strategies that enable quick problem identification and resolution. Drive Automation and Infrastructure as Code : Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines that enable reliable and consistent deployments. Create self-healing systems that can automatically respond to common failure scenarios. Establish SLOs and SLIs : Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring we maintain high reliability standards while balancing innovation speed. Incident Management and Response : Lead incident response efforts, conducting thorough post-morte
Jobiba hiring network
Senior Remote Remote Devops Engineer Observability Jobs
4,932 active opportunities · Updated for September 2026
Market range: $127K – $174K/yrFresh results
15 shown
Explore current senior remote remote devops engineer observability jobs. Use filters to narrow by work mode, employment type, experience and date posted.
As an Enterprise Sales Engineer, you will provide technical expertise through sales presentations, product demonstrations, and supporting technical evaluations (POVs). Sales Engineers help qualify and close opportunities with customers and partners and have a voice with the product team to help prioritize features based on input from customers, competitors, and partners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Partner with the Sales team to articulate the overall Datadog value proposition, vision and strategy to customers Own technical engagement with customers during the trial phase. Communicate Datadog’s value based on activities and work with customers on any identified issues or concerns to successful conclusion Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Deliver product and technical briefings / presentations to potential clients Maintain accurate notes and feedback in CRM regarding customer input both wins and losses Proactively engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscape Who You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customer Someone with strong written and oral communication skills. This role requires an ability to understand and articulate both the business benefits (value proposition) and technical advantages of our offering Experienced in programming/scripting with any of the following: Java, Python, Ruby, Go, Node.JS, PHP, and .NET etc. Someone with a minimum of 3+ years in a Sales Engineering or DevOps Engineering role Able to sit up to 4 hours, traveling to and fr
As an Enterprise Sales Engineer, you will provide technical expertise through sales presentations, product demonstrations, and supporting technical evaluations (POVs). Sales Engineers help qualify and close opportunities with customers and partners and have a voice with the product team to help prioritize features based on input from customers, competitors, and partners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Partner with the Sales team to articulate the overall Datadog value proposition, vision and strategy to customers Own technical engagement with customers during the trial phase. Communicate Datadog’s value based on activities and work with customers on any identified issues or concerns to successful conclusion Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Deliver product and technical briefings / presentations to potential clients Maintain accurate notes and feedback in CRM regarding customer input both wins and losses Proactively engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscape Who You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customer Someone with strong written and oral communication skills. This role requires an ability to understand and articulate both the business benefits (value proposition) and technical advantages of our offering Experienced in programming/scripting with any of the following: Java, Python, Ruby, Go, Node.JS, PHP, and .NET etc. Someone with a minimum of 3+ years in a Sales Engineering or DevOps Engineering role Able to sit up to 4 hours, traveling to and fr
As an Enterprise Sales Engineer, you will provide technical expertise through sales presentations, product demonstrations, and supporting technical evaluations (POVs). Sales Engineers help qualify and close opportunities with customers and partners and have a voice with the product team to help prioritize features based on input from customers, competitors, and partners. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Partner with the Sales team to articulate the overall Datadog value proposition, vision and strategy to customers Own technical engagement with customers during the trial phase. Communicate Datadog’s value based on activities and work with customers on any identified issues or concerns to successful conclusion Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Deliver product and technical briefings / presentations to potential clients Maintain accurate notes and feedback in CRM regarding customer input both wins and losses Proactively engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscape Who You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customer Someone with strong written and oral communication skills. This role requires an ability to understand and articulate both the business benefits (value proposition) and technical advantages of our offering Experienced in programming/scripting with any of the following: Java, Python, Ruby, Go, Node.JS, PHP, and .NET etc. Someone with a minimum of 3+ years in a Sales Engineering or DevOps Engineering role Able to sit up to 4 hours, traveling to and fr
As a Key Accounts Enterprise Sales Engineer, you will provide technical expertise through sales presentations, product demonstrations, and supporting technical evaluations (POVs). Sales Engineers help qualify and close opportunities with customers and partners and have a voice with the product team to help prioritize features based on input from customers, competitors, and partners. What You’ll Do: Partner with the Sales, Product/Engineering, Customer Success, Professional Services, and executive leadership to articulate the overall Datadog value proposition, vision and strategy to customers Lead technical strategy for both short-term and long-cycle pursuits by driving the end-to-end technical sales agenda for Fortune 100 and comparable prospects — from pipeline generation through purchase decision Technically close complex opportunities through advanced competitive knowledge, technical skill, and credibility Present at executive and technical forums, lead workshops, and translate technical detail to business impact for CIOs, SRE/Platform teams, security/compliance, and engineering leadership Maintain accurate notes and feedback in CRM regarding customer input both wins and losses Proactively engage and communicate with customers and Datadog business/technical teams regarding product feedback and competitive landscape Who You Are: Passionate about educating customers on observability risks that are meaningful to their business, and able to build and execute an evaluation plan with a customer Someone with strong written and oral communication skills. This role requires an ability to understand and articulate both the business benefits (value proposition) and technical advantages of our offering Experienced in programming/scripting with one or more languages (i.e. Python, Go, Java, etc.) and familiarity with DevOps practices such as CI/CD and IaC workflows Demonstrated ability to land into net new logo accounts and collaborate with prospects through long, multi-quarte
About the Role Amplitude's Cloud Platform team builds the systems that every Amplitude engineer relies on every day to ship code — and we're rebuilding them for the AI era. As a Senior Platform Engineer, you'll own medium-to-high-complexity platform projects end-to-end and help shape a platform where AI agents are first-class users alongside humans: kicking off deploys, opening pull requests against infrastructure, and triaging incidents, so a single engineer can get the throughput of a team. You'll partner with Staff engineers and product teams to make Kubernetes effortless across the engineering org, building self-service automation and scalable AWS infrastructure that lets product teams ship faster, safer, and with less cognitive load. If you're excited about building the systems that other engineers will rely on every day, this role is for you. Key Responsibilities Lead high-impact platform projects — design and ship capabilities that move the needle on developer experience, reliability, or security, and set the bar for quality, testing, and safe deployment practices. Build the AI-augmented platform. Design tooling and workflows that help engineers get more out of AI-assisted development — think infra primitives that are easy to reason about, automated review, and policy-as-code that keeps the guardrails strong as AI shifts how code gets written. Own Infrastructure-as-Code for Kubernetes, AWS, and GCP using Terraform, Helm, Kustomize, and emerging tooling — and make it consumable enough that an LLM can safely PR against it. Evolve our CI/CD backbone (Argo CD / Workflows / Rollouts, GitHub Actions) to make deploys faster, safer, and easier to reason about. Instrument and operate. Drive observability with Datadog and Amplitude, own dashboards and SLOs, and use the data to push reliability forward. Participate in on-call, lead incident response when needed, and turn postmortems into durable platform improvements. Reduce toil and tech debt with pragmatic remediation
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection. Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed. Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution
Location Details: Canada, Remote At GoDaddy, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) , and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team Contribute to the development of GoDaddy’s eCommerce and SSO infrastructure and Kubernetes systems on AWS. On a day-to-day basis you will be working on the team who designs, writes, tests and deploys the infrastructure and application management software for GoDaddy’s eCommerce applications. Expect to learn every day. What you'll get to do... Work as a polyglot engineer, writing and maintaining Infrastructure as code with frameworks/ ecosystems such as Java, Unix CLI, and NodeJS Build and operate infrastructure workflows and deployment pipelines using Kubernetes, Argo Workflows, Argo CD, and GitOps practices Design, build, and own services and APIs in Java, running on Kubernetes-based platforms across AWS and distributed systems Develop and support application and infrastructure delivery pipelines, enabling reliable releases of eComm, Auth and Infrastructure services Collaborate closely with other GoDaddy departments to help advance security and technical standards, maintain regulatory compliances while operating eComm & Auth platforms Your experience should include... 5+ years of strong backend software engineering experience in Java Hands-on experience with Kubernetes, including Helm, Kustomize, or equivalent tools to deploy and manage backend services Experience building and operating high-volume, mission-critical production systems on AWS with continuous deployment (CD) practices Strong experience with infrastructure as code, supporting backend applications and services Experience with observability a
We're Hiring Senior AI Engineer (Generative AI Azure) Remote Immediate Joiners Preferred Salary: Up to 12 LPA We are seeking an experienced Senior AI Engineer to design, develop, and deploy enterprise-grade Generative AI solutions on Microsoft Azure. The ideal candidate will have strong expertise in Large Language Models (LLMs), Agentic AI, Retrieval-Augmented Generation (RAG), Model Context Protocol (MCP), Azure AI Services, and modern AI engineering practices. Technical Summary • AI & LLMs: Azure OpenAI, GPT-4.x, OpenAI, Prompt Engineering, Prompt Chaining, Function Calling, Structured Outputs, Tool Calling, JSON Schema, Model Evaluation, Guardrails, Fine-tuning Concepts • Agentic AI: Multi-step Reasoning, Planning, Memory Management, Tool Orchestration, Multi-Agent Systems, Human-in-the-Loop Workflows, Reflection, Context Management, AI Observability • Frameworks: Semantic Kernel, LangChain, LangGraph, AutoGen, Azure AI Agent Service • RAG & MCP: Retrieval-Augmented Generation, Vector Search, Semantic Search, Hybrid Search, Embeddings, Knowledge Grounding, Citation Generation, Document Ingestion, Chunking, MCP Architecture, MCP Servers & Clients • Azure Technologies: Azure AI Foundry, Azure AI Search, Azure AI Document Intelligence, Azure Machine Learning, Azure Functions, API Management, Logic Apps, App Service, Container Apps, AKS, Azure Storage, Data Lake Gen2, Azure Key Vault, Azure Entra ID, Azure Monitor, Application Insights, Event Grid, Service Bus • Programming & DevOps: Python, C#, REST APIs, FastAPI, ASP.NET Core, JSON, YAML, Git, Azure DevOps, Docker, Kubernetes • Data Platforms: SQL Server, Azure SQL, PostgreSQL, Cosmos DB, Snowflake, Microsoft Fabric, Azure Databricks, Delta Lake, Vector Databases (Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant) • Security & Governance: Responsible AI, Prompt Injection Prevention, RBAC, Content Filtering, Data Privacy, GDPR, ISO 27001, Azure Key Vault, Audit Logging & Monitoring • Nice to Have: Microsoft Copilo
Location Details: Remote, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team... The Securities Analytics and Products Group is responsible for developing and maintaining sophisticated software solutions to safeguard GoDaddy's ecosystem. We are seeking a dedicated Senior Software Engineer with a strong background in software development and a keen focus on security. The ideal candidate will have proven hands-on experience in software development, with a consistent track record of working with the latest software technologies while prioritising security best practices. As a key member of our engineering team, you will play a crucial role in designing, developing, and implementing secure software solutions to protect our organisation from cyber threats. You will get to work with some of the brightest minds to build secure, highly available, fault-tolerant, and globally performant microservices-based platform deployed on the AWS cloud, using the newest technology stack. While the role is primarily backend-focused, you'll also get opportunities to contribute to frontend features as needed, giving you exposure across the full stack. What you'll get to do... Design, develop, and maintain secure, highly available, fault-tolerant, and globally performant code deployed on AWS cloud. Ensure code quality through extensive unit and integration testing Own frontend features and UI components across projects on an as-required cadence, from design through delivery Investigate and resolve production issues, ensuring your team's DevOps on-call responsibilities Contribute to the technical documentation, code reviews,
Location: Remote, India At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our team As a Senior Manager of Site Reliability Engineering, you will lead GoDaddy's SRE Center of Excellence in India. You will hold a senior role in the overall engineering leadership team. You will support GoDaddy's growth of its engineering footprint in India and develop a high-performing, cross-time-zone, globally unified SRE team. The object
About Runway Runway is a collaborative business planning platform designed to make business intuitively understandable for everyone. Our Mission : To make business accessible and understandable to everyone. We believe that teams that understand the “why” behind their work are more productive and make better decisions. True alignment and collaboration come from having a shared source of truth that everyone understands. Our Approach : Runway replaces traditional spreadsheets with a modern planning platform that brings clarity and context to business operations for all teams — not just finance. Just as Figma made design accessible across the organization, Runway does the same for business planning. Why It Matters: Understanding requires more than just access to numbers; real collaboration happens when teams see how their work fits into the bigger picture. By providing this context, Runway helps teams save time and move faster. Our Customers : World-class companies like AngelList, Superhuman, Stability.AI, ConvertKit, Lambda Labs, Lob, and SandboxVR rely on Runway to run their businesses more efficiently. Our Investors : We are supported by a select group of investors that we admire, including Garry Tan (YC & Initialized), a16z, Elad Gil, Naval Ravikant, Dylan Field (founder of Figma), Eric Ries, Claire Hughes Johnson (COO of Stripe), Henry Ward (founder of Carta), Akshay Kothari (COO of Notion), Eugene Wei, Lenny Rachitsky, Nikita Bier, Scott Belsky, Soleio Cuervo, Balaji Srinivasan, and many others. Working at Runway We're remote-first, so you can work from anywhere in North America. We strive to be clear in our communication and over-communicate by default. It’s early in our journey, so you'll have an opportunity to shape not just our product, but the company itself: how we work together, what makes us stand out, and who we hire. Below are the values we share as a team — if these resonate with you, you may enjoy working here. (And if you don't like them, please t
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team... We are seeking a highly skilled Senior Security Engineer to join our advanced Security Operations team. This role is focused on leading complex incident response and forensic investigations across Windows, macOS, Linux, and AWS environments while helping modernize security operations through automation and AI-driven capabilities. The ideal candidate is a hands-on security expert with deep AWS security expertise, strong threat detection and digital forensic skills, and experience leveraging AI and machine learning technologies to improve detection, response, and operational efficiency. You will play a key role in protecting critical assets, conducting high-impact investigations, mentoring team members, and driving the evolution of our security program against sophisticated and emerging threats. What You'll Get to Do... Lead high-priority incident response and forensic investigations, serving as the primary escalation point for advanced analysis, containment, recovery, root cause determination, and executive-level reporting. Drive threat detection and response across AWS, Windows, macOS, Linux, and endpoint security platforms, leveraging services such as GuardDuty, Security Hub, Detective, CloudTrail, IAM, VPC Flow Logs, and SentinelOne. Conduct malware analysis, host and cloud forensics, evidence collection, and threat hunting activi
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers and Architecture teams, your primary focus will be on ensuring production systems remain operational at all times, while continually setting and achieving long-term performance, reliability and scalability goals in a platform with an exponential growth plan for the coming years. With Okta’s increased dedication to ensuring customer availability expectations are exceeded in every way, you will play a key role as we evolve our system architecture to meet the demands of enormous growth and support the hundreds of millions of users who rely on us to provide uninterrupted access to business-critical enterprise and consumer applications. Skills Exceptional communication skills, including technical writing in the English language Systematic problem-solving approach, coupled with a strong sense of ownership and drive Understanding of microservices, cloud infrastructure (AWS, Azure), databases (SQL, No-SQL, Key/Value), containers (docker, kubernetes), web technologies (web sockets, http) and networking (SSL, routing, VPN) Live and breathe SLIs, SLOs, error budgets and SLAs Strong belief in automating everything and reducing toil for yourself and teammates Loves to work as a team, but is able to work effectively in a remote environment where tasks may be self-driven Knowledge
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role: Sentry is expanding our Technical Customer Success team to support our growing customer base and drive deeper adoption of the Sentry Platform worldwide. As a Technical Customer Success Manager (TCSM), you’ll be a Sentry product expert responsible for ensuring customers are successfully onboarded, achieve maximum value from our platform, and uncover new opportunities for growth through additional use cases and products. You’ll play a key role in helping customers realize measurable outcomes with Sentry. In this highly cross-functional role, you’ll collaborate closely with Account Executives (AEs), Sales Engineers (SEs), and Engineering teams to ensure customers’ technical and business goals are met. This role requires strong technical expertise and a deep understanding of the Software Development Life Cycle (SDLC) and related technologies. If you’re a technologist with experience supporting technical products in customer-facing roles—and you’re eager to join a fast-growing team delivering real value through an exceptional product—we’d love to meet you. In this role you will: Become a Sentry product expert and support customers by understanding their needs and helping them achieve their goals using the Sentry Platform. Drive customer success and health through effective onboarding, adoption, value realization, and retention. Collaborate with customer key stakeholders to define business and technical objectives, and work with the customer team to achieve them. Act as a trusted, strategic advisor to each assigned customer—driving best practices, innovation, and long-term success. Partner closely with the sales te
Get new senior remote remote devops engineer observability jobs by email
Daily job updates · Unsubscribe anytime