About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo
Jobiba hiring network
Cloud Operations Engineer Jobs
2,329 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary The Executive Director, Digital Engineering- Aetna Member Care and Journey Services is a senior technology leader responsible for setting the technical vision, architectural direction, and engineering execution for member centric services. This role leads large-scale engineering teams that build high-performance backend APIs, microservices, and cloud-native systems that power member experiences across digital, agent, provider, and partner channels. The leader ensures exceptional service stability, resiliency, innovation velocity, and alignment with enterprise user experience and operational goals. Key Responsibilities 1. Backend API & Microservices Engineering Leadership • Lead the design, development, and delivery of scalable backend systems, APIs, and microservices powering member-facing capabilities. • Define API contract standards, and integration patterns used across Member Services platforms. • Drive service modernization by adopting cloud‑native architectures, containerization, and event-driven patterns. 2. Service Stability, Observability & Resiliency • Establish standards for availability, resiliency, performance, and disaster recovery across all services. • Implement SLO/SLI/error budget frameworks, health checks, and high‑availability architectures. <p
About Us: Paytm is India's leading mobile payments and financial services distribution company. Pioneer of the mobile QR payments revolution in India, Paytm builds technologies that help small businesses with payments and commerce. Paytm’s mission is to serve half a billion Indians and bring them to the mainstream economy with the help of technology. About the role: As an Senior Engineering Manager, you will be developing the detailed design structure, implementing the best practices and coding standards, leading a team of developers for successful delivery of the project. You will be working on design, architecture and hands-on coding. Requirements: 9 to 13 years in Technical development with 5+ years in Providing technical leadership for high performance teams. Work closely with business and product teams to understand the requirements, drive design, architecture and influence the choice of technology to deliver solutions working closely with architects and leadership team. Build robust, scalable, highly available and reliable systems using Micro Services Architecture based on Java, Spring boot. Improve Engineering and Operational Excellence by identifying and building the right solutions for observability and manageability. Keep the tech stack current with the goal to optimize for scale, cost and performance. Migrate workloads to public cloud. Attitude to thrive in a fun, fast-paced environment. Serve as a thought leader and mentor on technical, architectural, design and related issues. Proactively identify architectural weaknesses and recommend appropriate solutions. Preferred Qualification : Bachelor's/Master's Degree in Computer Science or equivalent Skills that will help you succeed in this role: Tech Stack: Lang: Java, DB: RDBMS, Messaging: Kafka/RabbitMQ, Caching: Redis/Aerospike, Micro services, AWS. Strong experience in scaling, performance tuning & optimization at both API and storage layers Hands-on leader, and problem solver with a passion for e
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Cloud Support Engineer — AI/ML & Programmability (Night Shift) Location: Pune, India Snowflake's Support team is expanding. We're looking for a Cloud Support Engineer who enjoys working with data and solving a wide variety of problems, drawing on hands-on experience across operating systems, database technologies, big data, data integration, connectors, and networking. Our mission is to make Snowflake the preferred platform for running all AI, ML, data science, and data engineering workloads. You'll join a highly productive, fast-moving team supporting Snowflake Cortex and our ML product lines — work that is central to delivering on Snowflake's AI Data Cloud mission. Snowflake Support is committed to providing high-quality resolutions that help customers deliver data-driven business insights and results. We are a team of subject matter experts working collectively toward our customers' success, building partnerships by listening, learning, and connecting. Snowflake's values shape how we deliver world-class Support: putting customers first, acting with integrity, owning initiative and accountability, and getting it done. As a Cloud Support Engineer, you'll be the technical partner our customers turn to for guidance on using Snowflake effectively. You'll also be the voice
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Health100 is an AI‑native health technology platform that unifies pharmacies, providers, insurers, PBMs, and digital health solutions into a single, consumer‑focused ecosystem. Powered by Google Cloud AI, we’re reimagining personalized and connected health experiences. As a Senior Software Engineer for Health100, you will play a crucial role within a collaborative team — designing, developing, and maintaining backend services and APIs while ensuring releases are well-coordinated, fully prepared, and successfully deployed to production. The ideal candidate brings strong technical expertise in modern backend development, excellent problem-solving skills, and a proactive approach to production monitoring, issue triage, and cross-team coordination. This position is critical in maintaining high engineering standards, ensuring smooth release cycles, and driving operational excellence across the development lifecycle. *This role can be based anywhere in the US; hybrid or remote with preference for candidates to work out of our corporate headquarters in Woonsocket, RI. Responsibilities: Partner with technical leaders and the open-source community to contribute to technical designs, frameworks, roadmap definition, and requirements-gathering. Provide domain knowledge and engineering insight to guide early designs, ac
We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. Step Into the Multiverse: Your Next Adventure Starts Here At Wizards of the Coast, we harness the power of imagination and connection to create unforgettable experiences. We create entertainment that inspires creativity, sparks passion, forges friendships, and fosters communities around the globe. In every pursuit our mission is to inspire a lifetime love of games. Whether it's through the strategic depth of Magic: The Gathering®, the rich storytelling of Dungeons & Dragons®, or our AAA digital game studios, we build worlds that bring people together, spark creativity, and fuel adventure. As we continue to grow and explore new realms, we're seeking passionate, curious, and innovative minds to join the adventure. Do you have the versatility to operate at the intersection of game development, cloud infrastructure, and platform reliability? We are seeking a Senior Cloud Platform Engineer to join our Central Technology Build Engineering team as we develop and curate an Unreal Engine ecosystem to share across our internal game studios. You will join a specialized team responsible for unblocking developers, ensuring build stability, and optimizing iteration workflows across multiple concurrent titles. This is a hands-on IC role where engineering excellence meets operational reliability. You’ll collaborate closely with game development teams to ensure a seamless experience for all Unreal-powered studios under our banner. In this role, you will design, build, and run scalable AWS architectures that serve as the backbone
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake runs large scale cloud infrastructure to deliver its own service — production and internal deployments, Kubernetes fleets, CI/CD, etc. Our cloud spend is in billions of dollars per year. The Cloud Efficiency team builds a unified, self-serve cloud efficiency platform along with AI skills and agents that makes spend observable, attributable, governable while driving recommendations and optimization of our cloud spend. AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Design, develop, and maintain scalable platform for resource ownership registry, usage attribution, utilization measurement, and cost modeling. Build AI agents, tools and automation to enhance system monitoring, alerting, and root cause analysis. Improve and optimize data ingestion, storage, and query efficiency for cloud utilization, cost and efficiency data at scale. Collaborate with teams across Snowflake to understand attribution and observability needs and implement solutions that improve operational visibility. Contribute to open-source and industry best practices in monitoring and distributed systems monitoring. Ensure high availability, reliability, and performance of team-managed platforms by participating in on-call rotations and incident management. Partner with Finance, Product and Engineering
Location Details: India, Remote At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our Team Do you want to be an Information Security Lead at GoDaddy? GoDaddy’s Security organization is looking for a Cloud Security Engineer. We work out large-scale and cross-company security challenges while ensuring that partnership with the development and operational communities remains front of mind. At GoDaddy, Security Engineers apply their strong hands-on technical skills to craft scalable solutions for multiple problems. You must communicate with GoDaddy Engineering teams, perform security assessments, prioritize security risks, and design. We, as a team, implement high-quality security engineering solutions! What you'll get to do... The Senior Cloud Network Security Engineer will play a crucial role in designing, building, and securing large-scale, distributed cloud environments that support GoDaddy Services. This role operates at the intersection of cloud infrastructure, security architecture, and engineering execution. The successful candidate will collaborate closely with service teams, security leaders, and compliance partners to embed security-by-design principles into cloud services and internal platforms. This position requires advanced technical expertise in cloud-native security controls. It also needs a strong understanding of threat models in hyperscale environments. Additionally, it involves influencing architecture decisions across multiple teams. The role demands hands-on engineering, good judgment in ambiguous situations, and proficiency at translating security requirements into scalable, automated solutions. Bui
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Cloud Support Engineer (CSE) Job Description Snowflake seeks a Senior Cloud Support Engineers who combine technical expertise, customer empathy, and an AI-first mindset . You'll leverage and refine AI tools to accelerate troubleshooting, improve knowledge bases, and reduce time to resolution—safely and responsibly. Experience in 24x7 technical support, escalation handling, on-call rotations, and incident management is ideal. Key Responsibilities: As a Senior Cloud Support Engineer , you manage customer cases and are accountable for accelerating their time-to-resolution, ensuring platform stability, and driving a world-class support experience. Operating as a full-stack technical support resource, this role blends deep hands-on troubleshooting with the customer empathy and communication skills of a trusted advisor. Customer Value and Incident Ownership Own the Customer Experience: Manage customer issues from initial triage through resolution and follow-up, ensuring clear communication and timely updates. Deliver Support Value with AI: Leverage AI assistants and diagnostics to accelerate triage and root-cause analysis while maintaining accuracy and safety standards. Outcome-Based Support: Focus on business impact, ensuring resolutions fix issues, reduce recurrence, and
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote-first work , and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands. . Hiring and how we work We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. . See yourself at Twilio Join the team as Twilio’s next Software Engineer L2 to work in the Cloud Infrastructure. About the job This position is needed to evolve and maintain fundamental Compute infrastructure, collaborating with a passionate team to enhance the system capabilities. Key responsibilities include develop scalable cloud-native environments, VM orchestration, AWS-ASG Auto Scaling group of EC2 instances, hardened base AMIs, and secure container images while maintaining critical OS libraries. You'll implement solutions for providing robust legacy platform support and cutting-edge cloud-native technologies, all driven by automation and best practices. Join us in shaping the future of our Compute infrastructure Responsibilities In this role, you’ll: Collaborate with Tech Leaders, Architects and other Engineers to develop solutions for complex problems in distributed computing and infrastructure management. Automate solutions for operational issues,
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . Software Dev Principal Engineer - Cloud-Native Microservices & Network Security About the Role SonicWall is a global leader in cybersecurity, offering firewalls, secure access solutions, and threat intelligence that protect millions of networks for enterprises, governments, and service providers worldwide. SonicWall’s portfolio includes next-generation firewalls (NGFW), centralized management, wireless network management, secure mobile access, cloud security, and endpoint security. All these solutions are coordinated and managed through SonicWall's Unified Management (UM) architecture. We invite you to join our team in developing innovative network security management solutions for SonicWall's Network Security Manager (NSM), a centralized management and analytics platform within the SonicWall Unified Management (UM) ecosystem. In this role, you will help build secure and scalable software capabilities that integrate with SonicWall's Unified Management (UM) platform and Shared Services Platform (SSP) to meet the unique operational needs of enterprise customers, Managed Service Providers (MSPs), and Mana
SonicWall is a cybersecurity forerunner with more than 30 years of expertise and is recognized as a leading partner-first company, ensuring our partners and their customers are never alone in the fight against cybercrime. With the ability to build, scale and manage security across the cloud, hybrid and traditional environments in real-time, SonicWall provides relentless security against the most evasive cyberattacks across endless exposure points for increasingly remote, mobile and cloud-enabled users. With its own threat research center, SonicWall can quickly and economically provide purpose-built security solutions to enable any organization—enterprise, government agencies and SMBs—around the world. For more information, visit www.sonicwall.com or follow us on Twitter , LinkedIn , Facebook and Instagram . About the Role SonicWall is a global leader in cybersecurity, offering firewalls, secure access solutions, and threat intelligence that protect millions of networks for enterprises, governments, and service providers worldwide. SonicWall’s portfolio includes next-generation firewalls (NGFW), centralized management, wireless network management, secure mobile access, cloud security, and endpoint security. All these solutions are coordinated and managed through SonicWall's Unified Management (UM) architecture. We invite you to join our team in developing innovative network security management solutions for SonicWall's Network Security Manager (NSM), a centralized management and analytics platform within the SonicWall Unified Management (UM) ecosystem. In this role, you will help build secure and scalable software capabilities that integrate with SonicWall's Unified Management (UM) platform and Shared Services Platform (SSP) to meet the unique operational needs of enterprise customers, Managed Service Providers (MSPs), and Managed Security Service Providers (MSSPs) around the globe. You will design and build c
We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team. What you’ll be doing: Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc. Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities. Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests. Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites. Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins). Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters. Proven track record debugging large-scale, cloud-native stacks across ne
Get new cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime