Jobiba hiring network

Lead Cloud Infrastructure Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead cloud infrastructure engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

S
Sofi
📍 Seattle• Full-time
1mo ago

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The role We are looking for a Staff Cryptography Engineer to help lead SoFi’s enterprise readiness for post-quantum cryptography and long-term cryptographic resilience. This role sits within our Security Assurance organization and will partner closely with security, engineering, infrastructure, product, and business stakeholders to assess, modernize, and strengthen cryptographic controls across the enterprise. This is a highly cross-functional and strategic role. The ideal candidate brings deep applied cryptography expertise, strong engineering and architecture judgment, and the ability to drive complex security initiatives in ambiguous environments. You will help build SoFi’s cryptographic inventory, identify where cryptographic keys, protocols, algorithms, certificates, and libraries are used, and guide teams through the implementation of quantum-resistant and crypto-agile solutions. What you’ll do Lead efforts to assess and improve SoFi’s enterprise cryptographic posture, with a focus on post-quantum cryptography preparedness and crypto-agility. Build and maintain an inventory of cryptographic assets, including keys, certificates, algorithms, protocols, libraries, services, and business-critical systems. Partner with product, engineering, infrastructure, cloud, and security teams to identify cryptographi

restagileai
View job →
R
Roblox
📍 San Mateo• Full-time• From $326.1K/yr
1mo ago

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when

pythonjavaaws
View job →
D
1mo ago

Data Semantics is Datadog’s authority on semantic knowledge, providing shared infrastructure that powers both Datadog’s product experiences and AI capabilities. As Datadog continues its investment in OpenTelemetry-native observability, semantic interoperability, and AI-powered workflows, this team sits at the center of some of the company’s most strategic platform initiatives. As a Staff Engineer, you will serve as a technical leader for the team, balancing stewardship of critical production systems with the exploration of new platform capabilities that improve how telemetry is modeled, understood, and consumed across Datadog. You will work closely with engineering and product partners to define standards, drive technical direction, and deliver solutions that scale across Datadog’s observability platform. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the technical direction of the Data Semantics team while remaining deeply hands-on in design, implementation, and delivery. Build and scale semantic infrastructure that bridges OpenTelemetry, Datadog-native telemetry, cloud-provider telemetry, and customer-defined data models. Drive platform initiatives focused on schema evolution, telemetry standardization, data insights, and semantic interoperability across Datadog products. Partner with engineering and product teams across the platform to define standards, align stakeholders, and deliver high-leverage platform capabilities. Mentor engineers through design reviews, technical guidance, operational excellence, and long-term career development. Participate in on-call rotations and lead investigation and resolution efforts for complex production incidents affecting critical platform services. Who You Are: You have significant experience designing, opera

aigorust
View job →
M
Mongodb
📍 Ireland• Full-time
1mo ago

The data management software market is transforming how organisations build and run applications. MongoDB is the leading developer data platform and the first database provider to IPO in more than 20 years. Join us at the forefront of data and application development. MongoDB Technical Services Engineers combine deep technical expertise with exceptional problem-solving and customer-service skills. You’ll advise customers and resolve complex challenges across MongoDB Core, drivers, Atlas, Cloud Manager, cloud platforms, and infrastructure. We’re looking for candidates based in Dublin to join our vibrant office and collaborative in-office team. This is a five-day role with one of the following schedules: Tuesday–Saturday, Sunday–Thursday, or a five-day pattern covering both Saturday and Sunday. Under our hybrid model, employees on weekend schedules are expected to work from the office two days per week. Cool things you’ll do You’ll help customers troubleshoot complex issues and run critical MongoDB workloads with confidence. You’ll: Solve customer challenges across architecture, performance, recovery, and security Lead investigations from diagnosis to resolution, providing clear, actionable guidance Partner with Product Management and Engineering to advocate for customers and improve MongoDB Build tools, documentation, and training while mentoring peers and raising technical excellence What you need We value curiosity, adaptability, strong technical foundations, and a genuine desire to help customers. You should bring many of the following: 5–6 years of experience in technical support, systems engineering, database administration, SRE, or a related field Experience running complex, mission-critical production database systems Strong Linux and systems engineering skills, including performance, memory, I/O, storage, networking, security, clustering, and troubleshooting A solid understanding of networking concepts and protocols, including DNS, TCP/IP, and SSL/TLS Ability

javascriptpythonjava
View job →
M
Mongodb
📍 Dublin• Full-time
1mo ago

The worldwide data management software market is massive (IDC forecasts it to be $138 billion by 2026). At MongoDB, we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the center of innovation and creativity. MongoDB is seeking a Sr. Staff Software Engineer to join the Atlas Core Data Services organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product, along with the API Platform and Developer Tools. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Core Data Services organization builds the software that manages the Atlas cluster infrastructure hosted on the three major cloud providers (AWS, Azure, and GCP), as well as the software that manages the MongoDB database hosted on that infrastructure. We are constantly challenged to design features that ensure Atlas clusters are secure, available, durable, and performant while running large-scale, critical workloads. The Sr. Staff Engineer in this role will drive innovation across the organization and the company, setting technical standards and direction that enable future growth and velocity. We are looking for engineers with the experience and high standards needed to lead at that scale. Our organization champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies systems expertise to build the foundational infrastructure of a popular database, join us. Let's build a faster, more reliable, and highly scalable database platform together. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Define standards and vision for the mission-critical Atlas SaaS data

javamongodbaws
View job →
O
Okta
📍 Tel Aviv• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Identity Security Posture Management (ISPM) Team At ISPM (formerly Spera Security), we are revolutionizing identity security for modern organizations. Our platform provides security leaders with end-to-end contextual visibility into their identity landscape—reducing risk, enhancing operational efficiency, and ensuring continuous compliance. We’re a fast-moving, high-impact team looking for an exceptional Staff Software Engineer to help shape the future of identity security and platform architecture. Hybrid work - 3 days a week from the office. What You’ll Be Doing As a Staff Software Engineer, you will serve as a technical authority and play a pivotal role in scaling our core platform. Your focus will span architecture, security research, and engineering excellence: Lead & Architect: Design, build, and scale high-performance software while driving architectural best practices across the R&D organization. Security & AI Innovation: Drive deep-dive research into cloud, SaaS ecosystems, and AI agents to build proactive detection models, mitigation strategies, and threat intelligence. Engineering Excellence: Elevate code quality, testing standards, system monitoring, and documentation across teams. Cross-Functional Ownership: Partner closely with Product and R&D to translate complex security requirements into scalable technical solutions. Scale & Optimize: Identify and resolve systemic performance bottlenecks and architectural challenges.

pythonawsrest
View job →
O
Okta
📍 India• Full-time
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Required Skills: 9+ years in Python and its libraries (e.g., Pandas, Boto3) for data manipulation, ETL processes, and developing serverless functions to interact with foundational AI services on platforms like AWS Bedrock . 6+ years with modern front-end frameworks (e.g., React, Next.js, TypeScript), and an ability to collaborate effectively with Product & Design. Deep hands-on experience with AWS solutioning , including designing and deploying production applications using core services such as AWS EC2, AWS Lambda, Amazon S3, Amazon RDS, and API Gateway . Familiarity with Infrastructure as Code (IaC) tools like CloudFormation or Terraform is essential. Proven experience building and deploying production-grade AI applications and services. Experience with agentic frameworks like LangChain, LangGraph, and LangSmith is a strong plus. Strong background in building distributed systems, microservices, and resilient APIs (REST/GraphQL) on cloud platforms (AWS/GCP/Azure). Demonstrated ability to influence technical strategy and lead architecture decisions across global teams. Excellent communication skills with experience working effectively across time zones and cultures. Working experience with foundational AI services, with a desire to use a unified platform like AWS Bedrock for Generative AI development and deployment Education and Certifications A Bachelor’s degree in Computer Science, Information Systems, or equivalent years of industry experience AWS cl

typescriptpythonreact
View job →
C
Coinbase
📍 - USA• Full-time• Remote• From $186.1K/yr
1mo ago

Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Senior Software Engineer on the AI Platform team within the Platform group, you'll build and operate the LLM and agent infrastructure that every team at Coinbase depends on. This team owns the company's single path to large language models and the full agent lifecycle: build, deploy, run, observe, and improve. You'll lead multi-quarter technical initiatives across the platform, from gateway and runtime systems to knowledge bases and applied AI agents, directly shaping how Coinbase scales AI across the organization. What you'll do: Own the architecture and delivery of core platform systems including the LLM Gateway (60+ models, auth, PII redaction, fallbacks, cost optimization), AI Hub, and agent runtime with microVM sandboxes and governed MCP gateway Drive the design and implementation of Knowledge Base infrastructure, connecting data sources to auto-provisioned vector and markdown stores queryable by any agent Lead AI FinOps capabilities including spend attribution, governance, and cost optimization across all AI workloads company-wide Partner across engineering, security, legal, finance, product, and external partners at frontier labs and major cloud providers to ship high-impact platform capabilities Build evaluation and observability tooling including LLM-as-judge harnesses, full tracing, and feedback loops that let subject matter experts refine production a

REMOTEpythonawsmicroservices
View job →

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infr

javavueaws
View job →

Overview The Associate Director of Service Engineering leads the reliability, availability, and operational excellence of Natera’s lab-facing platforms. This role ensures that clinical systems, laboratory equipment workflows, and data pipelines operate with high reliability, scalability, and compliance in a regulated healthcare environment. You will lead a team responsible for production stability, incident response, and service health, partnering closely with Production Engineering, Lab Operations, Bioinformatics, Infrastructure, Facilities, and Compliance to support mission-critical genetic testing and diagnostics. Key Responsibilities Leadership & Team Development Lead, mentor, and scale a team of Service Engineers / SREs supporting clinical production systems Establish clear expectations around ownership, on-call readiness, and operational excellence Drive hiring, onboarding, performance management, and career growth Foster a blameless, learning-oriented culture focused on patient impact and reliability Service Reliability & Production Operations Own reliability and availability for production services supporting laboratory operations, reporting, and customer delivery Define and manage SLAs, SLOs, and operational KPIs aligned with clinical and business priorities Lead major incident response, ensuring rapid triage, clear communication, and thorough post-incident reviews Oversee on-call rotations, escalation paths, and operational playbooks Ensure operational readiness and go-live support for new assays, pipelines, and platform capabilities Technical Strategy & Execution Partner with Engineering and Development teams to design resilient, fault-tolerant systems Drive best practices for monitoring, alerting, logging, and observability across lab and cloud platforms Reduce operational toil through automation, tooling, and process improvements Advocate for reliability, performance, and scalability requirements early in t

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leader in full stack automation fabric solutions for mission-critical business processes. With the first SaaS-based composable automation platform specifically built for ERP, we believe in the transformative power of automation. Our unparalleled solutions empower you to orchestrate, manage and monitor your workflows across any application, service or server — in the cloud or on premises — with confidence and control. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for a Principal Engineer, Products & Platforms to join our Product engineering team to provide technical leadership across Redwood’s Workload Automation Platform, defining architecture, driving modernization, and influencing engineering strategy across multiple teams. You will be instrumental in the design, development, and enhancement of our platform, building high-quality, scalable, and secure software that powers enterprise data exchange for more than 1,000 customers worldwide. As a Principal Engineer, you will: Technical Leadership & Architecture: Define the path forward for complex engineering problems, establish best practices, lead design review and technology decisions, and mentor the team to excel, while driving the architecture, security, compliance, and observability of our Java/Spring Boot microservices. Platform and Infrastructure Ownership: Drive the deep understanding, architecture, and evolution of our core platform and infrastructure, focusing on resilience, communication between components, and scalability. Cross-Team Collaboration: Facilitate and drive cross-team collaboration with Product, QA, and other engineering groups to ensure end-to-end alignment and successful product deliv

javaawskubernetes
View job →
RS
12 days ago

OUR MISSION At Redwood, we empower our customers with lights-out automation for their mission-critical business processes. ABOUT US Redwood Software is the leader in full stack automation fabric solutions for mission-critical business processes. With the first SaaS-based composable automation platform specifically built for ERP, we believe in the transformative power of automation. Our unparalleled solutions empower you to orchestrate, manage and monitor your workflows across any application, service or server — in the cloud or on premises — with confidence and control. CORE VALUES One Team. One Redwood Make Your Own Weather Obsess over Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other YOUR IMPACT We are looking for an Engineering Manager, Products & Platforms to join our Product engineering team to lead a high-performing software engineering team while driving technical strategy across Redwood’s Workload Automation Platform. You will be instrumental in mentoring engineers, managing team deliverables, and guiding the design, development, and enhancement of scalable, secure software that powers enterprise data exchange for more than 1,000 customers worldwide. As an Engineering Manager, you will: People Leadership & Mentorship: Manage, coach, and grow a team of talented software engineers, supporting career development, conducting performance reviews, and fostering an inclusive, collaborative team culture. Technical Strategy & Architecture: Provide hands-on technical guidance, participate in design reviews, and define technical roadmaps for Java/Spring Boot microservices while ensuring high standards for architecture, security, and observability. Delivery & Platform Ownership: Oversee team execution, Sprint planning, and delivery timelines to ensure resilient, scalable core platform features and infrastructure. AI Integration: Research and apply AI/ML concepts and their usage to innovate and enhance the MFT (Managed

javaawskubernetes
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Establish and scale a new engineering organization focused on critical Exa platform services, ensuring a foundation of long-term success, technical excellence, and a high-performing culture. Own the successful delivery of complex, high-scale engineering features for the Exa platform, ensuring world-class security, reliability, and availability across multi-array and hybrid-cloud deployments. Define the technical vision and execution roadmap in close partnership with Product Management and Architecture, translating customer needs into a measurable business impact for Pure Storage. Drive a culture of operational rigor, owning the refinement of engineering processes around observability, CI/CD, and incident response, while actively mentoring the next generation of technical leads. WHAT YOU BRING Leadership and Scaling: Pro

awsci/cdrest
View job →
E
Everpure
📍 Bengaluru• Full-time
16 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE Join the Exa team and lead the charge in redefining enterprise storage by unifying block, file, and object protocols across hybrid-cloud environments. You will combine deep technical expertise in distributed systems with hands-on people leadership to guide architectural decisions and mentor high-impact engineers. This is a unique opportunity to build new engineering teams from the ground up and drive industry-leading innovation alongside Product and Architecture partners. Your work will directly impact how customers consume, scale, and operate mission-critical storage infrastructure. WHAT YOU'LL DO Drive End-to-End System Architecture: Lead the architectural evolution and end-to-end delivery of high-performance, resilient storage systems from initial design concepts to high-quality shipped products. Optimize for Modern Data Workloads: Design and implement robust algorithms and concurrent platform solutions engineered for modern data pipelines, AI infrastructure, distributed computing, and enterprise analytics. Resolve Complex System Engineering Challenges: Apply deep root-cause analysis and system-level insight to solve multi-threaded, high-concurrency performance and reliability issues across Linux platform internals. Cross-Functional Ownership & Leadership: Collaborate across product management, validation, and support teams to align technical roadmaps, establish architectural standards, and drive enterprise

pythonjavaaws
View job →
AC
16 days ago

Here at Appian, our values of Intensity and Excellence define who we are. We set high standards and live up to them, ensuring that everything we do is done with care and quality. We approach every challenge with ambition and commitment, holding ourselves and each other accountable to achieve the best results. When you join Appian, you’ll be part of a passionate team dedicated to accomplishing hard things, together. Summary: As a Principal Software Engineer at Appian, you will be the primary technical strategist, responsible for shaping the architectural foundation of our platform. Your role involves anticipating future challenges and implementing innovative solutions today. You will drive complex cross-functional initiatives, ensuring Appian remains a leader in the low-code and automation industry. Responsibilities: Define the long-term architectural vision and governance for the Appian platform. Identify systemic technical risks and lead task forces to resolve architectural bottlenecks. Develop internal tools and SDKs to simplify infrastructure and enhance developer productivity. Conduct deep-dive troubleshooting for complex production issues. Promote AI-native engineering practices and integrate AI features into the platform. Lead the development of core platform libraries and high-risk prototypes. Participate in the Architectural Guild and review high-impact design documents. Mentor lead and senior engineers, and represent Appian in the tech community. Ensure operational resilience with self-healing and highly available systems. Required Qualifications: Bachelor’s or Master’s degree in Computer Science, Information Technology, or related field. 15+ years of software engineering experience, with significant experience in architecting large-scale distributed systems. Strong understanding of data structures, algorithms, and design patterns. Proven transformational leadership in technological migrations or strategies. Expertise in Java and Cloud-Native ecosyst

javasqlaws
View job →
🔔

Get new lead cloud infrastructure engineer jobs by email

Daily job updates · Unsubscribe anytime