Jobiba hiring network

Cloud Operations Lead Jobs

15 active opportunities · Updated for September 2026

Fresh results

15 shown

Explore current cloud operations lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.

W
WPP
📍 ChennaiFull-time
4 days ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: The role is responsible for leading and overseeing end-to-end cloud operations, ensuring the availability, reliability, security, performance, and resilience of cloud platforms and services. It manages service monitoring, incident response, and incident resolution for production applications and cloud infrastructure, while ensuring operations teams are skilled and enabled to execute cloud-related requests with speed and diligence. The role also provides governance over a large third-party managed services organisation, ensuring delivery against agreed KPIs, SLAs, and operational commitments through effective service reviews, metrics reporting, performance management, and continuous service improvement. What you'll be doing: Responsible for overseeing cloud operations, driving operational excellence, and improving the operational landscape through automation and AI-driven solutions delivered by internal resources and third-party partnerships. Product: Work with product and engineering teams to define operational support patterns for each cloud product. Collaborate with business, archi

awsazuregcp
View job →

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

awsazuregcp
View job →
O
OpenAI
📍 San FranciscoFull-time
1mo ago

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

awslinuxrest
View job →
M
1mo ago

Cloud Operations Engineers are responsible for building internal tools and process automation. Day-to-day duties are creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases, and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace. We are looking for a Lead with strong technical leadership experience as well as technical depth who is looking to collaborate closely with Cloud Operations Engineering Management in building and maintaining a high-performing team that delivers high quality outcomes while fostering psychological safety and professional growth. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Core responsibilities Team leadership: partner with and assist COE Management with the tasks of providing ongoing technical feedback to engineers, support their growth and creating an inclusive team environment Execution and delivery: play a key role in guiding team members through project deliverables ensuring high quality outcomes while also assisting in meeting or resetting timelines when required Time management: between assisting team members with day to day tasks ranging from incident to project management Cross-functional collaboration: work closely with Product, Technical Services and R&D to surface team’s pain points and drive alignment with the goal of providing an excellent user experience to the end customer Coordinate with Lead counterparts within Cloud Operations as well as Technical Services to ensure our uptime guarantees to the MongoDB Atlas customer base Assist and collaborate with the team on scoping, designing, deploying and ongoing maintenance of systems that focus on reducing mean time to resolve customer incidents Detec

typescriptpythonjava
View job →
R
Replit
📍 Foster CityFull-time
1mo ago

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,

pythonawsazure
View job →
W
WPP
📍 ChennaiFull-time
4 days ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: WPP Media is embarking on a transformative data journey through Databridge —our internal comprehensive data strategy and framework designed to unify our fragmented global data landscape. We are moving from isolated, market-specific implementations to a standardized, scalable, cloud-agnostic, and AI-ready data platform built on a modern tech stack: dbt, Databricks, Python (dlthub), and GitHub Actions . As Data Operations Lead , you will establish and head our new Data Integration & Operations team in Chennai. This is a build-and-run role : you'll define how the team operates while leading day-to-day delivery of operational excellence across global data products. You will be the custodian of production —owning the operational layer including CI/CD automation, deployment pipelines, monitoring, data quality enforcement, incident management, and support. This role requires deep technical knowledge of Databricks and modern data platforms, alongside the ability to lead, mentor, and scale a growing team. What you'll be doin

pythonsqlazure
View job →
O
OpenAI
📍 WashingtonFull-timeRemote
12 days ago

About the team OpenAI’s mission is to ensure that artificial general intelligence benefits all of humanity. OpenAI for Government helps public-sector institutions use advanced AI responsibly to improve services and advance critical missions. The Partnerships team builds the external ecosystem and internal operating mechanisms required to deliver that work at scale. Role Summary We are seeking a Partnerships Operations Lead to build and run the operating foundation for OpenAI for Government’s Partnerships team. This person will serve as the operational connective tissue across the partner team, working closely with Account Directors, Solutions Architects, Revenue Operations, Deal Desk, and other cross-functional stakeholders to move opportunities from initial engagement through agreement and execution. They will help ensure a smooth handoff to Technical Success across accounts by establishing clear readiness criteria, capturing the right commercial and customer context, and aligning teams on partner roles, commitments, dependencies, and success outcomes. The ideal candidate is a true utility player who combines the relationship instincts of a growth oriented builder with the rigor of an exceptional business operator. They thrive in ambiguity, create processes that do not yet exist, and move fluidly between deal execution, workflow design, analysis, and cross-functional coordination. Key Responsibilities Own the end-to-end operating cadence for partner-sourced and partner-influenced opportunities, from intake and qualification through approvals, contracting, launch, and expansion. Build and continuously improve operating processes, standard procedures, stage gates, decision rights, and service expectations that make partner execution faster and more predictable. Operationalize new and existing government partner motions—including cloud providers, resellers, and sell-through models—by translating requirements that differ from standard commercial motions into clear work

REMOTEawsrestai
View job →
M
Modal
📍 New YorkFull-time
1mo ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We're hiring a Compute Strategy and Operations lead to own how Modal plans for and acquires GPU and CPU capacity. You'll size our infrastructure needs ahead of demand, source supply across hyperscalers, neoclouds, and datacenter operators, and negotiate and close the contracts to secure it. The compute you secure directly determines what Modal can sell and build. In this role, you will: Own end-to-end procurement of GPU and CPU capacity across hyperscalers, neoclouds, and datacenter operators Build and maintain a strong pipeline of supplier relationships Evaluate supply options on price, availability, hardware specs, networking capabilities, and SLA terms Negotiate and close contracts: reserved capacity agreements, spot arrangements, MSAs, DPAs, and order forms Work closely with our engineering teams to translate technical requirements into procurement specs Track

aigorust
View job →
CV
1mo ago

Role Summary We are seeking an experienced SOC / Security Operations Lead to oversee and strengthen the organization's Security Operations Center (SOC), threat detection, incident response, vulnerability management, data protection, and security monitoring capabilities. The ideal candidate will possess extensive experience in managing enterprise security operations, SIEM/SOAR platforms, threat hunting, incident response, DLP, endpoint security, and vulnerability management programs. The role requires leadership of a multi-functional security operations team responsible for protecting critical business systems, customer data, and digital assets while ensuring compliance with RBI regulations and industry security standards. Key Responsibilities Security Operations Center (SOC) Management -Lead and manage 24x7 Security Operations Center (SOC) functions. -Establish and enhance SOC processes, playbooks, escalation procedures, and operational metrics. -Ensure timely detection, triage, investigation, containment, and remediation of security incidents. -Develop SOC maturity roadmaps aligned with industry best practices and regulatory expectations. -Monitor security KPIs, SLAs, MTTR, MTTD, and incident response effectiveness. SIEM, XSIAM & Threat Detection -Lead implementation, administration, and optimization of: -Palo Alto Cortex XSIAM -SIEM Platforms -SOAR Platforms -UEBA Solutions -Threat Intelligence Platforms -Develop and tune correlation rules, detection logic, and analytics use cases. -Enhance detection coverage across cloud, endpoints, applications, networks, and third-party environments. -Drive threat hunting and proactive security monitoring initiatives. Incident Response & Threat Management -Lead enterprise cyber incident response activities. -Develop and maintain incident response plans, runbooks, and communication procedures. -Coordinate investigations involving malware, ransomware, phishing, insider threats, account compromise, fraud, and adv

CV
1mo ago

About Bazaarvoice At Bazaarvoice, we create smart shopping experiences. Through our expansive global network, product-passionate community & enterprise technology, we connect thousands of brands and retailers with billions of consumers. Our solutions enable brands to connect with consumers and collect valuable user-generated content, at an unprecedented scale. This content achieves global reach by leveraging our extensive and ever-expanding retail, social & search syndication network. And we make it easy for brands & retailers to gain valuable business insights from real-time consumer feedback with intuitive tools and dashboards. The result is smarter shopping: loyal customers, increased sales, and improved products. The problem we are trying to solve : Brands and retailers struggle to make real connections with consumers. It's a challenge to deliver trustworthy and inspiring content in the moments that matter most during the discovery and purchase cycle. The result? Time and money spent on content that doesn't attract new consumers, convert them, or earn their long-term loyalty. Our brand promise : closing the gap between brands and consumers. Founded in 2005, Bazaarvoice is headquartered in Austin, Texas with offices in North America, Europe, Asia and Australia. It’s official: Bazaarvoice is a Great Place to Work in the US , Australia, India, Lithuania, France, Germany and the UK! We are looking for a Senior Offensive Security Engineer who is a "Jack of all trades, Master of some." You don’t need to be a world-class Red Teamer AND a Cloud Architect; rather, you need a solid foundation in offensive security, operational experience in cloud platforms (AWS/GCP) and the maturity to manage our external penetration testing and bug bounty programs. As a leader within our proactive security strategy, you will drive the identification of complex vulnerabilities through expert management of third-party tests and by leading sophisticated, in-depth internal assess

pythonawsgcp
View job →

SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Role: Staff NOC Analyst (5 - 8 years) Location: Bangalore (24/7 Shift Environment) Role Summary We are looking for a Cloud Operations & Staff NOC Analyst who will act as the first line of operational defense for enterprise infrastructure, cloud platforms, and applications. This role requires strong real-time monitoring, incident response, and troubleshooting capabilities, along with a proactive mindset toward improving operational processes and reducing alert noise. Key Responsibilities Monitoring & Incident Management Monitor infrastructure, applications, and cloud platforms using tools such as New Relic, Datadog, Prometheus/Grafana, AWS CloudWatch, or GCP Monitoring Perform real-time alert triage, validation, and troubleshooting to restore services quickly Act as the first responder for incidents , ensuring minimal downtime and impact Identify false positives and reduce alert noise through analysis and tuning Incident Handling & Escalation Own and manage high-priority incidents (P1/P2), including: Driving incident bridges Coordinating with cross-functional

awsgcprest
View job →
CV
Company via Lever
📍 Pune, MaharashtraFull-timeHybrid
1mo ago

Perforce is a community of collaborative experts, problem solvers, and possibility seekers who believe work should be both challenging and fun. We are proud to inspire creativity, foster belonging, support collaboration, and encourage wellness. At Perforce, you’ll work with and learn from some of the best and brightest in business. Before you know it, you’ll be in the middle of a rewarding career at a company headed in one direction: upward. With a global footprint spanning more than 80 countries and including over 75% of the Fortune 100, Perforce Software, Inc. is trusted by the world’s leading brands to deliver solutions for the toughest challenges. The best run DevOps teams in the world choose Perforce. Priten Nayak, the VP of Cloud Operations at Perforce, is searching for a Senior DevOps Engineer III, India to design and build the next-generation cloud platform for Perforce’s SaaS product portfolio to ensure the security, reliability, and high availability of all our production & CI/CD environments and applications. In this vital role, you will drive the design, development, & Implementation of automated tools and technologies to enable efficient delivery and service management of the production services & release deliveries. Drive the adoption of AI-assisted cloudops practices and agentic automation to improve operational efficiency, security, incident response, and platform reliability across the software delivery lifecycle.

ci/cdairust
View job →
M
1mo ago

As our Accounts Receivable Manager, Marketplace, you will lead the financial operations for our rapidly expanding cloud marketplace ecosystem (AWS, Azure, and GCP). This is a strategic backfill for a key leadership role within our Global Order-to-Cash organization. You will own the daily financial operations of MongoDB’s cloud marketplace business across AWS, Azure, and GCP. Ensure seamless integration between partner portals and internal finance systems while safeguarding transaction accuracy, revenue integrity, and operational continuity. Proactively identify risks and implement scalable solutions as transaction volumes grow. Based in our Gurugram office, you will sit at the intersection of Finance, Sales Operations, and Cloud Engineering. You aren't just managing a ledger; you are managing the financial health of our most critical growth channel, ensuring that high-volume marketplace transactions are billed accurately and collected with surgical precision. What You’ll Do (Day-to-Day) Marketplace Operations Leadership: Serve as the functional successor in managing MongoDB’s cloud marketplace presence. Oversee the daily flow of transactions through AWS, Azure, and GCP, ensuring seamless integration between partner portals and internal finance systems Accounts Receivable Supervision: Maintain rigorous oversight of the AR lifecycle for the marketplace portfolio. This includes supervising the accuracy of invoices, monitoring aging reports, and ensuring that payouts are reconciled daily against recognized revenue Lead end-to-end AR lifecycle management for the marketplace portfolio: Oversee invoice accuracy and completeness to maintain reliable financial records, monitor aging accounts and drive a proactive collections strategy to optimize cash flow, meet/exceed assigned cash targets on a monthly and quarterly basis & participate actively in month-end and close activities, including reporting. Cash Application & Reconciliation Excellence: Ensure timely and accu

mongodbawsazure
View job →
W
WPP
📍 LisbonFull-time
4 days ago

WPP is the trusted growth partner for the world’s leading brands. We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth. We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise. Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow. For more information, visit WPP.com. Why we're hiring: At WPP, technology is at the heart of everything we do, and it is the Technology Operations teams mission, as part of Enterprise Technology , to enable our stakeholders to collaborate, create and thrive. Enterprise Technology is undergoing a significant transformation to modernise ways of working, shift to cloud and micro-service-based architectures, drive automation, digitise colleague and client experiences and deliver insight from WPP’s petabytes of data. This role will carry out the effective and efficient everyday technology operations for WPP ET. A trusted pair of hands to deal with level 1 and 2 issues as they present to the IT Service Desk and a trusted resource for Infrastructure and Management personnel to assist with project work when needed. The role will report into the Enterprise Technology Operations Lead and work closely with other teams within Enterprise Technology. What you'll be doing: Deliver world class, on-site support services to WPP employees, agencies, and visiting clients, operating within predefined structur

M
Mongodb
📍 JapanFull-time
7 days ago

MongoDB Pre-Sales Solutions Architects are technical business advisors who help customers design, justify, and adopt reliable, scalable systems using MongoDB’s data platform. They own the technical strategy across complex opportunities—from discovery and qualification through architecture, proof of value, executive alignment, and successful adoption—and connect technical decisions to measurable business outcomes. You’ll partner closely with Account Executives, Sales Leadership, Customer Success, Professional Services, and ecosystem partners to shape multi-threaded account strategies, build champions, de-risk complex architectures, and drive expansion. You’ll serve as a trusted advisor to developers, architects, operations leaders, and business executives, helping organizations modernize legacy systems, build AI-powered applications, and realize measurable value from MongoDB. . We are looking to speak to candidates who are based in Tokyo for our hybrid working model. As an ideal candidate, you will have: Ideally 8 to 11 years of related experience in a customer facing role, with 5 to 7 years of experience in pre-sales with enterprise software Minimum of 3 years experience with modern scripting languages (e.g. Python, Node.js, SQL) and/or popular programming languages (e.g. C/C++, Java, C#) in a professional capacity Experience designing with scalable and highly available distributed systems in the cloud and on-prem Demonstrated ability to lead architecture reviews for complex, multi-component applications and platforms, identifying risks, evaluating trade-offs, and providing clear guidance to modernize, optimize, and de-risk the solution Excellent presentation, communication, and interpersonal skills, with the ability to convey complex technical and business concepts in a clear and compelling manner to technology and business leadership Ability to partner with Sales Leadership and Account Executives on multi-threaded account and territory strategies, prioritize oppor

pythonjavanode.js
View job →
🔔

Get new cloud operations lead jobs by email

Daily job updates · Unsubscribe anytime