Jobiba hiring network

Lead Cloud Operations Engineer Jobs

6,876 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current lead cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

M
1mo ago

Cloud Operations Engineers are responsible for building internal tools and process automation. Day-to-day duties are creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases, and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. We are looking for a Lead with strong technical leadership experience as well as technical depth who is looking to collaborate closely with Cloud Operations Engineering Management in building and maintaining a high-performing team that delivers high quality outcomes while fostering psychological safety and professional growth. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Core responsibilities Team leadership: partner with and assist COE Management with the tasks of providing ongoing technical feedback to engineers, support their growth and creating an inclusive team environment Execution and delivery: play a key role in guiding team members through project deliverables ensuring high quality outcomes while also assisting in meeting or resetting timelines when required Time management: between assisting team members with day to day tasks ranging from incident to project management Cross-functional collaboration: work closely with Product, Technical Services and R&D to surface team’s pain points and drive alignment with the goal of providing an excellent user experience to the end customer Coordinate with Lead counterparts within Cloud Operations as well as Technical Services to ensure our uptime guarantees to the MongoDB Atlas customer base Assist and collaborate with the team on scoping, designing, deploying and ongoing maintenance of systems that focus on reducing mean time to resolve customer incidents Detec

typescriptpythonjava
View job →
W
WPP
📍 Chennai• Full-time
15 days ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The role is responsible for leading and overseeing end-to-end cloud operations, ensuring the availability, reliability, security, performance, and resilience of cloud platforms and services. It manages service monitoring, incident response, and incident resolution for production applications and cloud infrastructure, while ensuring operations teams are skilled and enabled to execute cloud-related requests with speed and diligence. The role also provides governance over a large third-party managed services organisation, ensuring delivery against agreed KPIs, SLAs, and operational commitments through effective service reviews, metrics reporting, performance management, and continuous service improvement. What you'll be doing: Responsible for overseeing cloud operations, driving operational excellence, and improving the operational landscape through automation and AI-driven solutions delivered by internal resources and third-party partnerships. Product: Work with product and engineering teams to define operational support patterns for each cloud product. Collaborate with business, archi

awsazuregcp
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

awsazuregcp
View job →
E
15 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE As a Staff Security Operations Engineer , you will own and continuously mature capabilities across Attack Surface Management, Vulnerability Management, Zero Trust, Secrets and Credential Security, Detection Engineering, and Security Automation . This is a hands-on role requiring strong security engineering expertise combined with the ability to build teams, establish operating processes, measure outcomes, and drive remediation across Engineering, Cloud, Infrastructure, and Product organizations. You will partner closely with GISO leadership and global security teams to translate security strategy into measurable execution and risk reduction. WHAT YOU’LL DO Lead and mature enterprise Attack Surface and Vulnerability Management capabilities across cloud, infrastructure, endpoints, applications, and internet-facing environments. Drive risk-based vulnerability prioritization using asset criticality, exposure, exploitability, known exploitation, threat intelligence, and compensating controls. Establish operating processes, remediation SLAs, KPIs/KRIs, dashboards, and governance to measure and drive security risk reduction. Identify systemic security gaps and develop scalable technical and operational solutions. Provide technical leadership across Zscaler/Zero Trust, secrets and credential security, SIEM/detection engineering, EDR, cloud security, and security automation. Drive automation and integrations using APIs, sc

awsazuregcp
View job →
A
Airbnb
📍 Remote - United States• Full-time• Remote
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: BizTech fosters culture and connection at Airbnb by providing reliable corporate tools, innovative products, and technical support for all teams. We drive technical breakthroughs and strategies that redefine what it means to belong anywhere, delivering greater value for the business and our people. The Global Operations team at BizTech manages production services across Airbnb’s corporate environment, delivering reliable operations through Observability, Incident Management, Core Operations, and AI-enabled automation. We partner across BizTech to scale service quality, efficiency, and resilience. The Difference You Will Make: As a Senior Staff Engineer in Operations, you will lead and mentor a high-performing team to scale our AI-enabled operations model and deliver AIOps solutions that streamline operational workstreams and help BizTech teams focus on their core work with confidence. Ops owns triage and resolution, proactive monitoring across networks, systems, applications, and cloud services via a homegrown observability platform, and drives process excellence through automation and shift-left programs. You will set the technical bar, model operational excellence, and ensure high-quality, reliable service. Your scope includes leading projects ac

REMOTEpythonawsci/cd
View job →

About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Order to Cash (OTC) team oversees the complete flow of commercial transactions from order intake and provisioning through billing, collections, and cash application — ensuring accuracy, compliance, and operational excellence in support of OpenAI’s mission to ensure artificial general intelligence benefits all of humanity. About the Role We are looking for a senior, hands-on operator to own Order Management and Billing execution across OpenAI’s cloud marketplace and partner ecosystem, including platforms such as AWS, GCP, Oracle Cloud, GovCloud, and future channels. This senior individual contributor role will translate partner requirements into scalable workflows and ensure launch readiness, accurate billing, and reliable daily execution. As a senior individual contributor within the Cloud Marketplaces team, you will own the end-to-end order-to-invoice lifecycle for your assigned portfolio. You will ensure that private offers, commercial terms, provisioning, pricing, usage, billing data, credits, settlements, and partner-specific reporting flow through our systems accurately, on time, and with audit-ready controls. You will implement and continuously improve the common cloud marketplace operating model, lead cross-functional execution for your assigned portfolio, and surface risks, requirements, and improvement opportunities. You will partner across Revenue Systems, Product, Engineering, GTM, Finance, Partner Operations, and external marketplace stakeholders. This role is critical to building the operational backbone for OpenAI’s expansion across cloud marketplaces and government-cloud channels. You will combine deep operational judgment with process and control execution, automation, clear communication, and hands-on problem solving to improve billing reliability, partner experience, customer outcomes, and financial integrity at scale. This role

awsgcprest
View job →
P
Pagerduty
📍 Atlanta• Full-time• $201K – $303.6K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. Senior Product Manager for AI and Automation PagerDuty is redefining how modern engineering and operations teams work. PagerDuty’s Automation Platform includes Workflows, Actions, Connectors, and a growing agentic layer built on Skills and Tools. This is the backbone of how teams eliminate toil, respond to incidents autonomously, and ultimately enable AI-native SRE agents. As Senior Product Manager for AI and Automation, you will own product strategy and execution across our Operations Cloud SaaS and on-premises automation products and lead the roadmap for the agentic automation experience we’re building for autonomous SRE agents. This is a high-visibility, high impact role that sits at the intersection of developer tooling, enterprise operations, and frontier AI product design. You will report directly to the Senior Director of Product Management for the AI & Automation group and will partner tightly with engineering, design, GTM, and enterprise customers. Key Responsibilities Define and drive the multi-year roadmap for Workflows and Actions, covering both cloud-delivered SaaS and on-p

awskubernetesgit
View job →
P
Pagerduty
📍 Toronto• Full-time• $156K – $236K/yr
1mo ago

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization. Senior Product Manager for AI and Automation PagerDuty is redefining how modern engineering and operations teams work. PagerDuty’s Automation Platform includes Workflows, Actions, Connectors, and a growing agentic layer built on Skills and Tools. This is the backbone of how teams eliminate toil, respond to incidents autonomously, and ultimately enable AI-native SRE agents. As Senior Product Manager for AI and Automation, you will own product strategy and execution across our Operations Cloud SaaS and on-premises automation products and lead the roadmap for the agentic automation experience we’re building for autonomous SRE agents. This is a high-visibility, high impact role that sits at the intersection of developer tooling, enterprise operations, and frontier AI product design. You will report directly to the Senior Director of Product Management for the AI & Automation group and will partner tightly with engineering, design, GTM, and enterprise customers. Key Responsibilities Define and drive the multi-year roadmap for Workflows and Actions, covering both cloud-delivered SaaS and on-p

awskubernetesgit
View job →
A
Asana
📍 San Francisco• Full-time• $250K – $330K/yr
1mo ago

As the Head of Engineering Operations, you will be a key leadership partner to the CTO and the engineering leadership team, owning the operational backbone of our global R&D organization. In this role, you will serve as a critical bridge to other cross-functional groups, driving operational excellence, organizational effectiveness, and strategic initiatives. We are looking for a strategic leader who can translate company strategy into actionable engineering plans while building lightweight operating models that accelerate velocity and maintain high quality. You will empower our engineering organization to thrive and deliver exceptional impact at scale. This role is based in our San Francisco or Vancouver office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Translate overarching company strategy into actionable engineering and operating plans in close collaboration with cross-functional partners. Design and build lightweight operating models, frameworks, and processes to optimize engineering execution and delivery, integrating evolving AI agentic engineering workflows. Hire, mentor, and lead a high-performing, global team of engineering operations professionals and Technical Program Managers to steer complex horizontal programs. Oversee financial and resource governance, including budget management, program spend, headcount strategy, and vendor/tooling allocations across public cloud and LLMs. Define and own the engineering metrics program (KPIs covering delivery, quality, reliability, efficiency, and capacity) to deliver data-driven insights to leadership. Serve as a trusted proxy and connective tissue for the CTO and R&D

aigorust
View job →
RS
10 days ago

About the Role Redwood is scaling public cloud infrastructure and AI features across multiple product lines, and we need a FinOps Lead to bring rigor, visibility, and accountability to that spend. This is a senior individual-contributor role with the authority to drive cross-functional cost governance directly. This role owns the translation of raw cloud cost data and AI spent into the models, forecasts, and governance mechanisms that let engineering, product, and executive leadership make informed decisions. This is not a bill-monitoring role. You will build the cost attribution infrastructure that ties cloud/AI spend to specific product lines and, ultimately, to ROI involving architecture and engineering teams to make the right trade offs and decisions in line with the strategic roadmap. The role will report to the Senior Director, Product Engineering Operations, and work closely with Cloud Engineering, the CPO's org, and engineering leadership to create cost visibility and defensibility informing leadership on efficiency strategies. Your ability to be an effective communicator, collaborative team player, and analytical thinker will be keys to success in this role. Responsibilities Cost Visibility, Attribution & Optimization Own and continuously improve cost models that attribute AWS (and other public cloud) spend by product line, team, and environment Drive tagging governance and hygiene, define standards, audit compliance, and close attribution gaps that prevent accurate cost-per-product reporting Build shared cost allocation models to drive transparency into per-team cost drivers where there are prevalent savings plans, reserved instances and network charges Build toward feature-level cost attribution that connects infrastructure spend to product ROI, not just aggregate bill totals Create and manage optimization programs with achievable savings targets, including working across engineering and finance to rightsize, clean up and modernize infrastructur

sqlawskubernetes
View job →
B
Baseten
📍 San Francisco• Full-time
1mo ago

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and

awsgcpci/cd
View job →
S
Sofi
📍 San Francisco• Full-time
1mo ago

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. Role Description As the Director of Corporate Infrastructure, you will drive efforts to oversee the design, implementation, and operation of our corporate networks. This includes the IT Infrastructure DevOps, Security and SRE teams. You have a deep understanding of infrastructure as code, configuration as code & networking technologies, strong business acumen, lead through data and metrics, and operate with a high level of rigor and accountability. You are responsible for the reliability, scalability, sustainability, and efficiency of SoFi’s network infrastructure. As a member of the Corporate Infrastructure leadership team, you are directly accountable for the teams that are responsible for driving efforts to evolve and build our next-generation corporate network (to include infrastructure and wifi). This includes managing all aspects of our network - engineering and operations to improve user experience and performance, and also supporting the multi-terabit backbone network that interconnects edge PoPs, corporate offices, data centers, and cloud gateways. You will lead a diverse team through highly technical problems to achieve our overall strategic, operational, and financial goals. You will build and lead a high-performance team, displaying technical proficiency to support and scale the

awsrestai
View job →
R
Replit
📍 Foster City• Full-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,

pythonawsazure
View job →
DC
15 days ago

Join Delphi - Where Innovation meets transformation At Delphi, we believe in creating an environment where our people thrive. Our hybrid work model empowers you to choose where you work—whether it's from the office, your home, or a mix of both—so you can prioritize what matters most. We are committed to supporting your personal goals, family, and overall well-being while driving transformative results for our clients. We welcome exceptional talent from anywhere across the globe. Interviews and onboarding are conducted virtually, reflecting our digital-first mindset. Rooted in the region, we specialize in delivering tailored, impactful solutions in Data, Advanced Analytics and AI, Infrastructure, Cloud Security, and Application Modernization. Whether it’s enabling predictive analytics , transforming operations with automation, or driving customer engagement with intelligent platforms, we are the trusted partner for organizations ready to embrace a smarter, more efficient future. We are looking for a hands-on Senior QA Consultant to lead end-to-end testing of AI and Generative AI applications across enterprise environments. This role combines strong expertise in traditional QA engineering with modern AI evaluation and validation practices. The ideal candidate will drive quality assurance initiatives for RAG pipelines, multi-agent systems, OCR and Speech-to-Text solutions while ensuring production grade quality outcomes in regulated industries such as Finance, Healthcare, and Insurance. The successful candidate will lead a small QA pod, collaborate closely with engineering, AI/ML, product, and business teams, and act as the client-facing QA owner for enterprise AI engagements. This role requires strong technical leadership, automation expertise, AI evaluation capabilities, and excellent stakeholder management skills. Experience Requirements • 10–15 years of experie

pythonjavasql
View job →
🔔

Get new lead cloud operations engineer jobs by email

Daily job updates · Unsubscribe anytime