GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a Senior Backend Engineer in GitLab’s Platform Enablement organization, you will help make GitLab easier to deploy, validate, and operate across environments. This team is centered on two major areas: Cloud Native deployment guidance and ephemeral environments. You’ll help shape GitLab’s next generation of self-managed deployment guidance as Reference Architectures evolve toward a model based on deployment patterns, workload characterization, component requirements, topology guidance, and scaling principles. You’ll also help improve production-like ephemeral environments so teams can validate changes e
Jobiba hiring network
Cloud Operations Engineer Jobs
2,329 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cloud operations engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.
At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: The Senior Software Engineer is responsible for building the tools required to help organizations succeed with Apple. Lead others on the agile team to break down problems and apply the appropriate designs and practices to build Jamf products. Subject matter expert in various Jamf components and product offerings. Mentor and coach others while delivering new components and features with high quality and reliability. You may be required to work periodically at a Jamf office or collaborative work location with other Jamf employees in your area for certain events or moments that matter. What you can expect to do in this role : Break down customer problems into work you and the team can execute on. Independently complete tasks from start to finish with high quality. Ability to communicate technical concepts to stakeholders. Use your knowledge of Engineering best practices to ask the right questions, solve problems and build great software with a high level of quality. Produce designs for new and existing features. Clearly communicate technical concepts with others in the organization (Technical Communication, Support, Product and Cloud). Performs all job responsibilities in alignment with the core values, mission and purpose of the organization. Adheres to the highest moral, ethical and legal standards to deliver and environment that promotes respect, innovation and creativity
At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. This role is offered as a hybrid in Brno, Czech Republic. We are only able to accept applications for those based in the Czech Republic or who have sponsorship to live and work in the Czech Republic. What you'll do at Jamf : At Jamf, we empower people to be their best selves and do their best work. As a Software Engineer, you'll be part of a team building macOS extensions that power Jamf Protect - a cross-platform enterprise endpoint security solution that helps administrators protect their organization's devices and data. You'll work with real-time monitoring, custom threat detections, and Apple's endpoint security frameworks, all optimized to preserve the Apple user experience. What you can expect to do in this role: Break down customer problems into actionable work for you and the team. Own tasks end-to-end and deliver high-quality results. Leverage enterprise-grade tooling - including AI - to work smarter and move faster. Mentor teammates using your deep product knowledge. Ask the right questions, solve problems, and build great software using engineering best practices. Communicate technical concepts clearly across teams - Engineering, Support, Product, and Cloud. Know when to reach out - and find the right p
We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Own the reliability architecture of the platform across regions and cloud providers Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices Set operational standards for the team: on-call quality, incident response, SLO discipline Mentor and technically develop the SRE team Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Qualifications 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms Proficiency in Python, Go, or a similar programming language Understand workload isolati
MongoDB is looking for an outstanding person to join our newly created Forward Deployed Engineering team and take on a key role in our extended R&D organization. Forward Deployed Engineering is linking the work of teams engaged on application modernization roles with our Product and Engineering teams. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Many organizations have built up large estates of legacy applications. Lack of scalability and resilience, long development times, operating cost, and inability to run on cloud are common issues with these applications. To address these issues, organizations are engaging in large transformational Application Modernisation programs. MongoDB is recognized as the developer data platform of choice for transactional systems that provide the best scalability, resiliency and developer experience in the cloud as well as on premises. Organizations are continuously migrating workloads from these legacy applications to new platforms, often based on microservices, using MongoDB. Such transformations are time intensive and often risky. Tooling based on generative AI promises to accelerate these transformations in a way never seen before. Forward Deployed Engineering is responsible for exploring the possibilities of Generative AI technologies and providing invaluable feedback to MongoDB’s R&D teams to drive future capabilities of MongoDB and the MongoDB ecosystem. Application Modernization Engineers will work alongside project teams that are executing Application Modernisation projects with customers. The successful candidate will be responsible for evaluation, build, and applying tools in modernization projects, facilitating the usage of such tools and processes across the different project teams, identifying opportunities for tooling deployment, selecting potential 3rd party tools, contributing to the development of tooling prototypes and helping to shape the produ
MongoDB is looking for an outstanding person to join our newly created Forward Deployed Engineering team and take on a key role in our extended R&D organization. Forward Deployed Engineering is linking the work of teams engaged on application modernization roles with our Product and Engineering teams. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. Many organizations have built up large estates of legacy applications. Lack of scalability and resilience, long development times, operating cost, and inability to run on cloud are common issues with these applications. To address these issues, organizations are engaging in large transformational Application Modernisation programs. MongoDB is recognized as the developer data platform of choice for transactional systems that provide the best scalability, resiliency and developer experience in the cloud as well as on premises. Organizations are continuously migrating workloads from these legacy applications to new platforms, often based on microservices, using MongoDB. Such transformations are time intensive and often risky. Tooling based on generative AI promises to accelerate these transformations in a way never seen before. Forward Deployed Engineering is responsible for exploring the possibilities of Generative AI technologies and providing invaluable feedback to MongoDB’s R&D teams to drive future capabilities of MongoDB and the MongoDB ecosystem. Application Modernization Engineers will work alongside project teams that are executing Application Modernisation projects with customers. The successful candidate will be responsible for evaluation, build, and applying tools in modernization projects, facilitating the usage of such tools and processes across the different project teams, identifying opportunities for tooling deployment, selecting potential 3rd party tools, contributing to the development of tooling prototypes and helping to shape t
We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w
We are hiring a Senior Software Engineer to join our Server Security team. The Server Security team is a development-focused group within MongoDB's core engineering organization. Operating "close to the bottom of the stack," the team builds features that enable database users to secure their data globally. You will work on critical components including: Cryptography: Queryable Encryption , at-rest data encryption, and fundamental cryptographic principles. Identity & Access: Authentication and authorization systems, TLS, and X.509 certificate management Network Security: High-performance, low-latency networking protocols (PKI, Hashing, CRLs) System Integrity: Resilience, observability, and compliance assurance within a large-scale distributed database Our team champions a strong culture of inclusivity, diversity, and collaboration. If you want to work on a collaborative team that applies distributed systems fundamentals to deliver core features of a popular database, join us! Let’s change what’s possible for application developers, system architects, and database operators. The role As a Senior Engineer, you will apply distributed systems fundamentals to deliver core security features. You will be a leader in improving MongoDB's security posture by owning features and leading investigations into complex areas of the codebase. What you’ll do: Build and test new security features in a large, feature-rich C++ codebase Work across engineering, cloud services, and support teams to coordinate feature rollouts and changes Stand for code quality and security best practices, assisting fellow engineers in writing well-reasoned, secure code Use strong diagnostic intuition to solve thorny technical issues related to distributed systems, concurrency, and OS internals This role can be remote or hybrid anywhere in the USA or Canada. We will prioritize candidates who are already located in one of these countries. Candidate Profile We are looking for a highly technical engineer w
The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency requests around the globe, and comply with various data sovereignty requirements. The SRE Team’s mission is to build this increasingly complex infrastructure, while continually lowering the operational burden associated with it, and increasing our internal visibility into the health of the system. We are strong believers in infrastructure-as-code and self-healing systems. The SRE Team is fully integrated with all the other engineering teams, and the teams work closely together with a soft and traversable boundary between their areas of responsibility. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Design and build the infrastructure for a global cloud service that comprises hundreds of thousands of MongoDB clusters, processes a billion metrics per day, and replicates tens of billions of database writes to our backup service Design, implement, and troubleshoot the automation and monitoring of services that seamlessly spans the globe - including several cloud providers Become an expert in infrastructure performance, helping us optimize from the application level all the way through the firmware Build for resilience. Our goal is that nobody’s pager goes off, ever. Are we there yet? No. Are we really close? Very. While we work on that - participate in a weekly on-call rotation Improve our infrastructure capabilities, optimizing for cost, simplicity, and maintainability Requirements 3+ years of experience running a mission critical service at scale in a Linux environment Firm grasp of at least one modern programming language, beyond basic scripting Familiarity with web and network protocols and standards (HTTP, TLS, DNS, etc) Bachelor’s deg
The MongoDB Product Security organization is a diverse collection of individuals working together to scale MongoDB’s security, both security of the products themselves and the security features we offer to customers. The team is responsible for the MongoDB Database Server ( Community and Enterprise editions). The MongoDB Product Security organization works with software engineers to design, implement, and operate systems in a manner that protects customer data. It is a multidisciplinary team that covers product, software, cloud, infrastructure, and operational security concerns. The team does the following: Build a developer driven security program where there is tight integration with engineering artifacts, process, and tooling Use software architecture and coding patterns to reduce the impact of security issues Be security subject matter experts for our tech stack and products This role can be based out of any MongoDB U.S. office or remotely within the United States. Who You Are With a strong security engineering background, you’re looking for a role that gives you the freedom to increase MongoDB’s resonance with customers by strengthening our core database products. You’re passionate about solving hard security engineering problems while putting a strong emphasis on customer experience, leveraging your own significant experience. You enjoy collaborating with different teams to innovate and implement pragmatic solutions. Responsibilities You will take ownership, define strategy, and drive improvement for parts of our program such as fuzzing, threat modeling, secrets management, or container security You will support engineering teams when they respond to security issues, such as providing risk analysis and proof of concept Advocate for and contribute to complex security projects from inception through completion Drive architecture, patterns, and processes across Server Engineering that make security the easiest path Partner closely with engineering teams to
The Infrastructure Engineering team is responsible for building and maintaining a self-service internal development platform that enables MongoDB engineering teams to reliably deploy and operate their own production services and products. We work with numerous engineering teams across the company to understand their infrastructure requirements and development workflows, develop broadly applicable self-service platform services and tooling, continuously monitor how platform services are being utilized, and look for ways to improve developer productivity through automation and education. We are big open source enthusiasts and use a number of open source tools in our stack (contributing upstream whenever possible). Some of the tools we use regularly include Go, AWS, Kubernetes, Crossplane, Terraform, Helm, Drone, Prometheus, and Grafana. However, technology is nothing without a stellar team of engineers that are focused on doing high quality work and working as a team to solve complex distributed computing and platform engineering problems. This is where you come in! We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Our ideal candidate 2+ years of experience managing and mentoring a team of 3+ engineers Has 5+ years of experience owning the design and implementation of large software/infrastructure projects Has built and operated large-scale distributed systems in cloud providers (AWS strongly preferred) Has a strong backend programming background. Fluency in Go is strongly preferred; deep experience with another compiled or strongly-typed backend language is acceptable Pragmatic, detail-oriented, self-motivated, and understands the benefits of collaboration Strong experience operating production Kubernetes clusters, not just deployed to it Has practical experience defining and operating against SLI/SLOs for services they owned Strong experience with observability tooling: metrics, logging, traces, Prometheus, Grafana, OpenTe
About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system
About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil
About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal a dependable, scalable foundation for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between physical hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Security Engineer, Host Assurance to help build the trust foundations for bare-metal platforms across OpenAI’s global infrastructure. This is a deeply hands-on engineering role for a builder who can design, implement, and operate the core security infrastructure that establishes trust in hardware platforms before they are eligible to run workloads. Success in this role requires strong technical judgment, the ability to work comfortably at low levels of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—including novel hardware platforms and emerging deployment models– to make the secure path the easiest path. This role is well suited for engineers who enjo
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets
Get new cloud operations engineer jobs by email
Daily job updates · Unsubscribe anytime