Senior Software Engineer - Observability and Reliability About the Role We are growing the engineering team and looking for engineers who have the chops to build and deliver world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible. What You Will Be Doing Build observability tools and platforms, including: metrics, logging, distributed tracing, dashboarding, alerting, application performance management Build with modern tools and languages like Go, Open Telemetry and Kubernetes Participate in on-call rotation and ensure uptime of services Create runtime tools/processes that optimize cloud triaging and limit downtime Define best practices around making our systems and services measurable Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies. We expect successful candidates to be coding a majority of their time Qualifications We Need Strong Computer Science fundamentals 5+ years industry experience building and maintaining high-quality software, especially software other engineers use You apply a product mindset to infrastructure systems and feel accomplished enabling others Desire to be a great teammate and have fun at work Strong sense of craftsmanship, and a healthy academic curiosity Qualifications We Want (also, skills you’ll learn!) Experience building systems for data analytics Distributed systems monitoring and profiling skills Knowledge of cloud application security models Administered cloud service infrastructure (GCP, AWS, Azure) Startup experience Additional Job details Additional Job details The base salary range for this position is $170k - $240k annually. Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience. Base pay is one part of the Total Package that is provided to compensate and recognize e
Jobiba hiring network
Senior Cloud Infrastructure Engineer Jobs
7,292 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current senior cloud infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Role Sigma is transforming how businesses run by delivering a high performance platform on the modern data architecture. Hence, we are growing the engineering team and looking for engineers who are excited to solve challenging problems, deliver impactful capabilities throughout our stack to build world-class technology. You will be part of a talented team of engineers with a shared mission to make data easily accessible for all users. What You Will Be Doing Solve challenging problems that arise in providing an interactive experience on data warehouses for data exploration and analysis Build with modern tools and languages like Rust, Go, GraphQL, Node, and Kubernetes Build backend distributed services, new algorithms and modern API to support a cloud application Triage product or system issues and debug/track/resolve by analyzing the sources of issues Design and implement new software features to support our fast growing user base Collaborate with peers and stakeholders through design and code reviews to ensure best practices amongst available technologies Qualifications We Need 5+ years industry experience building and maintaining high-quality software Experience building and deploying robust and secure web applications in a continuous deployment environment Desire to be a great teammate and have fun at work without compromising ownership towards your work Strong sense of craftsmanship, and a healthy academic curiosity to solve challenges at sigma Strong Computer Science fundamentals Qualifications We Want (also, skills you’ll learn!) Data driven aptitude and its application to solve distributed system problems Data model design, and API development experience SQL query optimization and database internals Administered cloud service infrastructure (GCP, AWS, Azure) Prior experience working at high growth company solving technical problems to enable continued success Additional Job details The base salary range for this position is $170k - $240
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our Toronto or Vancouver offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design primitives of at least one of AWS, Azur
Overview: We are looking for a Senior Data Engineer with deep expertise in Lakehouse architecture, real-time data streaming, cloud data infrastructure, and microservices development on Azure Kubernetes Service (AKS). You will play a central role in designing and delivering next-generation data pipelines, BI solutions, AI/ML platforms, streaming APIs, and scalable microservices that power Guidepoint's research and analytics products. This is a high-impact, hands-on engineering role. You will work closely with data architects, data scientists, analysts, frontend engineers, QA, and DevOps teams to translate complex business requirements into scalable, reliable, and observable data systems. This is a Hybrid role from our Pune office. What You'll Do: Data Engineering & Lakehouse Design, build, and maintain ETL pipelines, data ingestion workflows, and table schemas on Azure Databricks to support BI, analytics, and AI/ML use cases Architect and optimize the Lakehouse using Delta Lake on Databricks, ensuring reliability, performance, and cost efficiency Build and support data pipelines from business applications such as Salesforce, NetSuite, and other enterprise systems Develop and maintain Knowledge Graph models, entity relationship structures, and NLP-based insight pipelines Maintain data governance, data privacy standards, and compliance best practices throughout the data lifecycle Perform root cause analysis on data and processes to identify opportunities for improvement Collaborate with data architects, scientists, and business consumers to populate and optimize the data warehouse for reporting and analytics Microservices & AKS Development Develop and support scalable web APIs and microservices using Python and Azure Platform Services Build new applications, services, and platforms; optimize existing solutions and refactor legacy components using modern, scalable architec
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The role We are seeking a Staff Vulnerability Management Engineer to lead the most complex technical work in SoFi’s Vulnerability Management program. You will design and build scalable systems that identify, enrich, prioritize, route, and track vulnerabilities across applications, cloud and infrastructure, containers, software supply chains, and specialized hardware or firmware surfaces. This is a hands-on engineering role with broad technical influence: you will write production code, make architecture decisions, establish vulnerability management standards, and improve how teams understand and reduce vulnerability risk. You will partner with Engineering, Infrastructure, SRE, Compliance, Legal, and business stakeholders to accelerate remediation while protecting engineering velocity and customer trust. You will also serve as a senior technical responder for embargoed disclosures and zero-day events, lead root-cause analysis for high-impact vulnerability incidents, and mentor engineers. The ideal candidate combines deep vulnerability management expertise with strong software engineering judgment, systems thinking, and a bias for durable, measurable outcomes. What you’ll do Lead high-complexity vulnerability management initiatives and make architecture decisions for assigned program areas, from detecti
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We're hiring a Product Partnerships Manager to own Replit's most important strategic partnerships end-to-end: from identifying the opportunity to shipping the outcome and measuring the impact. This is a product-focused role where you will work with our consumer technology partners, such as Stripe, Shopify, Google, and the broader connector ecosystem. You'll work at the intersection of product strategy, ecosystem thinking, and deal execution. You'll need to be a serious Replit power user. You'll speak credibly with engineers. You'll negotiate with senior partner stakeholders. And you'll be the person accountable for turning ambiguous ecosystem opportunities into shipped product and business outcomes. What You'll Do Develop a clear point of view on Replit's partner landscape and build a prioritized pipeline of high-leverage opportunities across consumer tooling, AI tooling, cloud and infrastructure, and payments technology partners. Lead partner conversations from early exploration through joint product thesis, business case, commercial terms, launch plan, and post-launch iteration. Partner with Product, Engineering, Partner Engineering, Legal, Finance, Marketing, and Sales to turn agreements into shipped outcomes. Work directly with Partner Engineers to scope integrations, demos, prototypes, reference apps, and partner enablement assets. Define success metrics before every launch activation: retained usage, apps created, deployments, revenue, partner-sourced users, and use them to decide when to scale, iterate, or sunset a partnership. Build lightweight operating systems: partner scorecards, launch checklists, partner roadmap tracking, and repeatable frameworks for evaluating new opportunities. Represent
NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and dedicated people on the planet working for us. If you're creative and autonomous, we want to hear from you! What you'll be doing: Developing software to enable efficient network design, deployment and day 2 management. Building product focused software solutions, used by internal and external customers. Helping us as we transform our workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies. Owning and driving integrations with various service APIs such as Cloud Service Providers, to automate creation of environments and auto populate data sources in turn. Building on open source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations. Streamlining deployment mechanisms and life cycle operations Developing modern service architectures around streaming data and event pipelines. Working with infrastructure domain experts on true, zero touch deployment solutions and utilizing best of breed high performance computing management solutions. Be a proactive problem solver, looking out for new opportunities to improve our services and customer experience. Communicate readily with your peers across the organization, b
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is the identity standard. The Okta Platform is an independent and neutral platform that securely connects the right people to the right technologies at the right time. We help organizations do two things - secure and manage their extended enterprise, and transform their users’ experiences. Okta's Core Engineering team is responsible for building and evolving shared infrastructure and services that lay the foundation for what other engineering teams build on. We're in charge of common shared services like search, cache, configuration management, frameworks for async job management, and email pipeline, to name a few. We're cloud native, where redundancy, multi-tenancy, scale, resource optimization and resiliency are first class citizens. With Okta's mantra of 'Always On!' there's never a dull moment. Our biggest asset is our team of passionate engineers and technically minded managers. We're looking for a staff level backend engineer to join a team of highly skilled and talented team players who're proud of what they own and deliver. Our elite team is fast, creative and flexible; with a weekly release cycle and individual ownership we expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company that is changing the cloud computing landscape forever. You will: Work with engineering teams to design, develop and deliver cloud based infrastructu
Identity has become a critical to an organizations security posture. This is true across their entire digital landscape include cloud, traditional on prem and emerging technologies like agentic ai. Britive is at the forefront of a modern approach to delivering identity security with the only modern privileged access management platform that provides unified Privileged Access Visibility, Dynamic Privilege Management and Secrets Governance across infrastructure, platforms & SaaS. Our patent-pending technology is deployed at companies of all sizes around the world, including Fortune 500 companies. We have repeatedly ranked among the hottest Cloud Security startups. Britive is founded by CyberSecurity industry veterans with a successful prior exit and is backed by top-tier VCs. About You You are a passionate Senior Software Engineer who wants to develop and scale our multi-tenant SaaS applications on the AWS platform. You have strong hands-on technical expertise in a variety of big data technologies. From day one, you must be able to hit the ground running and bring all your experience to the team to contribute in the building of a great product. Most importantly, you have a positive “can do” attitude and a passion for delivering technical solutions in a fast-paced startup environment. Your Impact Key Responsibilities: Responsible for design and development of a large-scale data ingestion and analytics. Design and develop data pipelines for real-time and batch data processing for disparate datasets. Develop systems which can store data in highly normalized fashion to allow correlation with other data sources. Collaborate with product management and engineering teams to design and integrate software, conduct code reviews, and troubleshoot product issues. Perform proof of concepts to identify best design options including usage of AWS services. Resear
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . As a Software Dev Senior Engineer , you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error budgets, and continuously reducing toil through engineering. Key Responsibilities: Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs ) for all critical services, partnering with product and engineering leadership. Own the error budget framework: track consumption, enforce error budget policies, and drive reliability investments when budgets are at risk. Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into pro
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Software Engineer for Infrastructure Security you will be a part of the Information Security organization and report to the Senior Manager of Infrastructure Security. You will help shape the future of Platform Security at Roblox. We work closely with Production IAM, Network Security, and Cloud Security at Roblox. You Will: Identify security gaps and threats in our cloud and on premise infrastructure, partnering with Governance Risk and Compliance teams to create standards and policies along the way. This will help Roblox meet regulatory and compliance requirements. Harden our infrastructure by introducing secure by default configurations, designs and guardrails for all developers at Roblox. Own and drive solutions that enable Roblox engineers to design, build, and use infrastructure securely at scale. Work closely with other InfoSec teams (AppSec, D&R, GRC, CorpSec, CloudSec, NetSec) and partner with engineering teams across Roblox, specifically the Infrastructure organization, to ensure the secure outcomes of security and product driven initiatives. You Have: 5+ years of experience writing code and/or relevant technical experience. Experience with
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. Senior Staff Security Engineer The Senior Staff Security Engineer owns current and target-state data architectures and reporting while also designing, implementing, and monitoring cloud (AWS/GCP) infrastructure security controls; deploying, securing, configuring, and operating SIEM and other security resources; identifying, triaging, and remediating infrastructure and web vulnerabilities; leading incident triage and external-researcher engagement; mentoring junior staff; and helping shape secure, scalable approaches for AI-enabled tooling, automation, and emerging product capabilities. Role summary and core responsibilities 6+ years of relevant experience Candidates must demonstrate proficiency in cloud security, network security, URL filtering, common security frameworks, and CVE lifecycle management Prac
The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing (including our public-facing edge and internal service mesh), and observability and alerting systems. The Fleet Management team provides the core runtime environment that empowers our developers to build and ship products to delight our customers. We manage the end-to-end lifecycle of our Kubernetes fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As our infrastructure scales to support new use cases and products, we are spearheading a migration from Terraform-based Infrastructure as Code (IaC) to an Operator-driven lifecycle management model. This role can be based out of our Austin, Boston, Los Angeles, New York City, Raleigh, or San Francisco offices, remotely in the United States region, or our European office in Dublin. Responsibilities Contribute to developing and maintaining a scalable and secure runtime environment on top of Kubernetes that supports product needs across MongoDB Provide internal support for our Kubernetes ecosystem, partnering with engineering teams to help them solve domain-specific problems Participate in a 24/7 on-call rotation to resolve critical issues Prioritize blameless post-mortems and dedicate engineering time to systemic fixes, ensuring you aren’t paged for the same issue twice You may be a good fit if you Have 6+ years of experience in software development and operating distributed systems Are proficient in Go, Python, or a similar language, with a strong commitment to code quality and testing practices (writing unit, integration, and E2E tests) Have deep experience using and extending containerization technologies, preferably Kubernetes Have a solid understanding
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is seeking an experienced Senior Adobe Experience Cloud Engineer with a deep understanding of Adobe’s tech stack to join our growing team.The position will play a crucial role in designing, developing, and maintaining solutions that leverage Adobe technologies to meet the unique business needs of our business partners. You must be collaborative and able to build trusted partnerships with extended teams. A successful candidate will have the ability to balance priorities and collaborate with cross functional teams while delivering within an agile delivery framework and supervising key performance indicators. Responsibilities Solution Design & Development: Lead the development and implementation of custom solutions and integrations within Adobe Experience Cloud, including Adobe Experience Cloud solutions, including Content Management, Assets, Multi-Site-Management, and Cloud manager Architecture & Scalability: Architect and build scalable, high-performance systems and applications that meet business requirements and technical specifications. Integration & Optimization: Develop and maintain integrations between Adobe Experience Cloud products and other internal or third-party systems. Optimize existing systems for performance and reliability. Collaboration: Work closely with product managers, solution architects, and other stakeholders to gather requirements, define project scopes, and deliver high-quality software solutions. Agile Development:
Get new senior cloud infrastructure engineer jobs by email
Daily job updates · Unsubscribe anytime