The Security Libraries team owns the customer-side integrations behind Datadog’s run-time security products — App & API Protection , Workload Protection , and Code Security — shipping and maintaining security capabilities across seven open-source language libraries ( .NET , Java , Go , Node.js , Python , Ruby , PHP ) and a set of HTTP proxy integrations (Envoy, NGINX, and HAProxy), running inside thousands of production clusters worldwide. Recent work spans exploit prevention (RASP), WAF detections, API Security, code security (IAST and SCA), and AI-assisted onboarding. As Engineering Manager, you’ll lead part of this polyglot team, setting the technical bar and team culture while driving the pace at which new detections and AI-assisted capabilities reach customers. This is a hands-on role: you’ll balance people leadership, product and roadmap ownership, and the operational health of code that runs in production at massive scale, with room to grow into more of Datadog’s security portfolio over time. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead, grow, and develop a team of roughly 4-8 library engineers — coaching, giving direct feedback, and empowering senior ICs as technical leaders Own team delivery and productivity: planning, milestones, reviews, and the on-call rotation Set product direction with Product Management and balance priorities across App & API Protection, Workload Protection, and Code Security Stay technically close to the work — apply strong technical judgment, contribute code where it matters most, and keep quality and architecture high Build strong relationships and drive alignment across the language teams and product, backend, and frontend partners Shape team identity and culture, and own accountability when problems oc
Jobiba hiring network
Cluster Chef Jobs
313 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster chef jobs. Use filters to narrow by work mode, employment type, experience and date posted.
MongoDB is seeking an Engineering Manager to join the Atlas Organization. The organization is responsible for building MongoDB Atlas, our database-as-a-service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Data Federation & Archiving team is an engineering team responsible for the Atlas capabilities that allow customers to move data from hot to cold storage and run federated queries over that data. The team builds Atlas Data Federation, a distributed query engine that lets users query data across Atlas Clusters and cloud object storage through a unified service. The team also builds Atlas Online Archive which allows customers to move data from Atlas Clusters into fully managed cloud object storage while preserving a seamless query experience across hot and cold datasets. We are forming a new Atlas Data Federation & Archiving team in the Dublin area. The Engineering Manager who fills this position will be pivotal in growing that team. We are looking to speak to candidates who are based in Cork and would like a hybrid or in-office working model. What you’ll do Lead a team of motivated individual contributors who are eager to learn and grow Contribute to the code, design, and architecture of the systems your team develops Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 6 years of professi
MongoDB is seeking an Engineering Manager to join the Atlas Organization. The organization is responsible for building MongoDB Atlas, our database-as-a-service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. This includes developing software to interface with the three major cloud providers (AWS, Azure, and GCP) in order to bring security, durability, availability, and performance to all deployments of MongoDB. The Atlas Data Federation & Archiving team is an engineering team responsible for the Atlas capabilities that allow customers to move data from hot to cold storage and run federated queries over that data. The team builds Atlas Data Federation, a distributed query engine that lets users query data across Atlas Clusters and cloud object storage through a unified service. The team also builds Atlas Online Archive which allows customers to move data from Atlas Clusters into fully managed cloud object storage while preserving a seamless query experience across hot and cold datasets. We are forming a new Atlas Data Federation & Archiving team in the Dublin area. The Engineering Manager who fills this position will be pivotal in growing that team. We are looking to speak to candidates who are based in Dublin and would like a hybrid or in-office working model. What you’ll do Lead a team of motivated individual contributors who are eager to learn and grow Contribute to the code, design, and architecture of the systems your team develops Work with stakeholders throughout MongoDB to build our roadmap and product offerings Work with customers and support engineers to fix issues and become part of our on-call rotation Collaborate with team members to develop our codebase, best practices, and design principles Foster an inclusive and respectful work environment according to MongoDB's Core Values We’re looking for someone who Has at least 6 years of profes
The MongoDB Cloud Services Team is a diverse group of contributors working together to help our users manage MongoDB at global scale. The Cloud Team is responsible for MongoDB Atlas: our database as a service offering, and fastest growing product, which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. The Backup Team delivers essential infrastructure to help our customers in their hour of need - providing the ability to quickly restore a massive, distributed database to any point in time at the click of a button. The Backup Team’s mission is to make MongoDB backup more reliable, faster, and also cheaper. This team is responsible for the Backup Agent (Go), the extensive server-side infrastructure (Java) which manages 100s of TB of data and processes billions of operations per day, and the user interface (Javascript) that customers use to manage their backups. Common project themes are performance, scaling, and ease of use. We are looking to speak to candidates who are based in New York for our hybrid working model. We're looking for someone who is Skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Fond of chasing down tough problems in a distributed systems environment Cool under pressure - has wrangled production crises, and secretly finds this a little fun Experienced with Linux, and able to correlate application performance problems with underlying hardware limits Comfortable working across the stack of a modern web application Always striving to expand their knowledge Curious, collaborative and intellectually honest Responsibilities Work closely with product teams, considering the user’s perspective while helping the team achieve success Collaborate with team members over best practices and core concepts Hold yourself accountable to your actions, maintaining the balance between accomplishing goals with research & development Own our
The Atlas API Experience (APIx) teams are part of MongoDB Atlas Data Services, a diverse group of individuals who develop the capabilities to run MongoDB globally (see MongoDB Atlas ). Our software and services allow users to deploy fault-tolerant, scalable, globally distributed MongoDB clusters in minutes. APIx’s mission is to create delightful experiences that bring developers along on the journey from eager beginner to sophisticated MongoDB expert! As part of our team in APIx, you will be responsible for improving & extending our API platform & downstream tooling, such as the Atlas CLI . Our mission is to help Atlas customers automate their workloads easily through APIs & support internal teams contributing to the API. We are looking for passionate, intrinsically-motivated software engineers who lead by example, raise the bar for those around them, and want to make a broad impact. No prior experience with MongoDB technologies is required! During interviews, you will meet most of our team and have the opportunity to ask questions about working at MongoDB. We pride ourselves on our team's culture and on being an inclusive and collaborative group that Builds Together . This role can be based in our Dublin office or remotely within Ireland. What will you do? Propose, design, implement and support product features for the MongoDB Atlas API Platform & MongoDB Atlas DevTools, such as AtlasCLI Build tooling that enables MongoDB users and developers to succeed, using programming languages such as Java, Go, Python & Javascript/Typescript Design and develop software integration components utilizing MongoDB’s technology (i.e. database, mobile, search, etc.) in larger contexts and integrate into partner frameworks or solutions Mentor and provide technical guidance to other engineers through code and design reviews, pairing, and knowledge sharing Lead sophisticated projects end to end, breaking large efforts into incrementally shippable deliverables Investi
The MongoDB Atlas team is a diverse group of contributors working together to help our users manage MongoDB at a global scale. We are responsible for MongoDB Atlas: our database-as-a-service offering and fastest-growing product, which allows users to deploy fault-tolerant, globally distributed MongoDB clusters in just minutes. We're seeking a Software Engineer to join the Atlas Identity and Access Management (IAM) team. The IAM team owns authentication and authorization for all of MongoDB Atlas, spanning from the underlying platform to user-facing features and products. IAM is both a platform and a product team, serving internal engineers as well as external customers. We enjoy the challenge of keeping MongoDB Atlas secure while providing a best-in-class user experience for users ranging from startup to enterprise. We are looking to speak to candidates who are based in New York City, NY for our hybrid working model. Role Responsibilities Lead medium-sized, full-stack projects from technical design to delivery Collaborate with colleagues at all stages of the project lifecycle (ideation, requirements gathering, design, execution, and delivery) Own our leadership principles , and exemplify them in your work Continue to learn and be curious in growing your career! Candidate Profile 2+ years of professional experience building full-stack web applications Proficient in a modern compiled programming language (Java, Go, C#, C++, etc.) Willingness to learn JavaScript and/or TypeScript along with modern frontend technologies (React, Redux, etc.); prior experience a plus Excellent communication skills, both written and verbal Is collaborative, empathetic, and intellectually honest Success Measures In 1 month, you’ll have shipped code into production In 3 months, you’ll have collaborated and delivered on a project with other engineers on the team In 6 months, you’ll have led the technical design, execution, and delivery of a project About MongoDB MongoDB is built for change, em
The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency requests around the globe, and comply with various data sovereignty requirements. The SRE Team’s mission is to build this increasingly complex infrastructure, while continually lowering the operational burden associated with it, and increasing our internal visibility into the health of the system. We are strong believers in infrastructure-as-code and self-healing systems. The SRE Team is fully integrated with all the other engineering teams, and the teams work closely together with a soft and traversable boundary between their areas of responsibility. We are looking to speak to candidates who are based in New York City for our hybrid working model. Responsibilities Design and build the infrastructure for a global cloud service that comprises hundreds of thousands of MongoDB clusters, processes a billion metrics per day, and replicates tens of billions of database writes to our backup service Design, implement, and troubleshoot the automation and monitoring of services that seamlessly spans the globe - including several cloud providers Become an expert in infrastructure performance, helping us optimize from the application level all the way through the firmware Build for resilience. Our goal is that nobody’s pager goes off, ever. Are we there yet? No. Are we really close? Very. While we work on that - participate in a weekly on-call rotation Improve our infrastructure capabilities, optimizing for cost, simplicity, and maintainability Requirements 3+ years of experience running a mission critical service at scale in a Linux environment Firm grasp of at least one modern programming language, beyond basic scripting Familiarity with web and network protocols and standards (HTTP, TLS, DNS, etc) Bachelor’s deg
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement b
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I
About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A
About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late
About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
About the Team The Applications Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably. In this role, you will: Design and build the development and production platforms that power our products, enabling reliability and security at scale Ensure our infrastructure can scale to the next order of magnitude Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed. You might thrive in this role if you: Have 5+ years building core infrastructure Have experience operating orchestration systems such as Kubernetes at scale Have experience building abstractions over cloud platforms Take pride in building and operating scalable, reliable, secure systems Are comfortable with ambiguity and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and
Get new cluster chef jobs by email
Daily job updates · Unsubscribe anytime