Atlas Growth is a cloud engineering group whose mission is to guide customers through their app development journey—from cluster configuration, data modeling and load testing, to running a production workload at scale. We use an in-house experiments platform which helps us validate our features quickly, releasing only the work that positively impacts our customers. Our engineers participate in cross-functional “squads” with product, design, analytics, and research focusing on a single metric (e.g. retention). Our engineering team is part of a larger Atlas Core Engineering org, building foundational elements of MongoDB’s developer data platform. Atlas Growth 2 builds customer-facing features in Atlas and sits alongside other Growth engineering teams. Recent projects include an AI Chatbot for cluster creation, a recommendation system that offers tips for better database performance, and a pricing page designed to optimize conversion rates. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Role Overview Atlas Growth seeks a mid-level software engineer (Software Engineer 3). SE3s are solid contributors to projects they work on and often lead projects of their own. They act in accordance with MongoDB’s core values and leadership principles, and are actively working toward a Senior role. Candidate Profile 3+ years of software engineering experience, with fluency in TypeScript/JavaScript, and experience with a modern framework (e.g. React) Proficiency in Java, Go, C++/C, or a similar compiled language is a plus Experience writing database queries, either document-based or relational Experience writing and reviewing technical specs, and leading small projects Interest in A/B testing or product design Expectations Contribute readable and well-tested code to ongoing projects Collaborate closely with product and design partners to implement and iterate on new customer facing features Write scope and technical spec docs for new projects
Jobiba hiring network
Cluster Chef Jobs
313 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current cluster chef jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the role (Remote) We're hiring a Senior Product Designer to join our global UX team. You'll own design within your product cluster — working directly with the engineers and PMs in those areas to define problems, shape solutions, and ship work that holds up. This is an end-to-end role. You lead discovery, define the interaction model, produce specs, and stay close through implementation. You're also expected to contribute to the systems and standards the whole team relies on. The Senior Product Designer will report to our Product Design Manager, and will work with our EU, UK and US based designers. You're part of a distributed team, so async clarity and proactive communication are as important as craft. What you’ll do Lead the design of complex AI-powered interactions in your product cluster, accounting for trust, transparency, fallback states, and human-in-the-loop considerations Embrace the use of modern AI-based tools like Claude, Figma Make, Cursor, Vercel, Loveable or equivalent not as theatre, but as a meaningful multiplier to your workflow. Partner with our UX research team to lead interviews, usability tests, and competitive analysis and translate findings into clear product direction and guidance Develop presentations, wireframes, mockups, and prototypes that effectively communicate interaction and design intent across web and mobile surfaces. Partner with PMs to shape requirements and with engineers to protect design intent through delivery Contribute to Showpad's design systems — adding components, documenting patterns, raising consistency across product areas Communicate design rationale through strategic storytelling: you can walk a PM, an engineer, and a senior stakeholder through the same decision and each one gets what they need Establish scalable AI-assisted workflows for your own work — research synthesis, rapid prototyping, concept generation — and help others on the team use them effectively Deliver UX solutions that measurably improve
About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Software Engineer on Spark Platform, you will execute across the surfaces of our in-house Spark deployment that serves the entire company. The work spans Spark runtime upgrades and performance, multi-tenant scheduling and executor bin-packing on Kubernetes, cluster lifecycle automation, and the observability and incident automation that keep the platform sustainable. You will move between layers as the work demands — picking up the next high-leverage problem regardless of where it sits — and partner closely with the rest of the team and with platform consumers across the company. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Build and operate an in-house Spark platform that runs at company-wide scale, spanning runtime, scheduler, reliability, and user-facing tooling. Drive multi-tenant scheduling, executor bin-packing, and cost-aware placement that let a small team serve dozens of consumer teams. Own pieces of cluster lifecycle automation — provisioning, upgrades, capacity changes, and node-failure handling — at a scale where these stop being manual events. Build the observability and incident automation that make the platform debuggable end-to-end and keep on-call sus
About the Team The Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company at significant scale and continue to expand the workloads, capabilities, and consumer base we serve. Orchestrating and operating thousands of Spark cluster deployments is a complex distributed system problem which the team invests heavily in runtime optimization, systems architecture, multi-tenant scheduling, and end-user tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the technical direction for our in-house Spark deployment and shape the architecture that will run DoorDash's data, analytics, and ML compute for the next five years and beyond. You will own the deep, cross-cutting problems that span the runtime, the shuffle service, the scheduler, and the overall service reliability — making the architectural calls that compound across the platform's lifetime. You will partner with the Engineering Manager on technical roadmap, hiring, and team shape, and act as the senior technical voice in cross-team partnerships with Data Engineering, ML Platform, and product engineering teams that depend on the platform. You must be located in San Francisco, Sunnyvale, Seattle, or New York City for this hybrid position. You will report into the Engineering Manager on our Spark Platform team. You're excited about this opportunity because you will… Set the multi-year technical direction for an in-house Spark-on-Kubernetes platform — runtime, shuffle, scheduler, reliability — and make the architectural calls that compound for years. Own the deepest distributed-systems problems on the team: shuffle architecture, multi-tenant scheduling, runtime performance, and the failure modes that only show up at scale. Partner with the Engineering Manager on technical roadmap, hiring, inte
Responsible for developing and implementing effective safety strategies, policies, and procedures for the Surguja Cluster. Their primary objective is to achieve Zero incidents by implementing robust Safety Management Systems (SMS) and Integrated Management Systems (IMS), while also ensuring compliance with industry standards. They oversee the comprehensive Occupational Health and Safety Management Systems (OHSMS) programs, manage safety initiatives, conduct audits, and coordinate emergency response plans to maintain a safe working environment and minimize risks in coal mining operations. Source: Adani Group | Job ID: 55009
The Sr. Officer will play a pivotal role in ensuring the safety and security of Adani Solar’s vertically integrated solar PV manufacturing ecosystem at Mundra. This role is critical to safeguarding the operational integrity of the 800-acre Mundra Electronic Manufacturing Cluster, which houses cutting-edge technologies and facilities for solar PV manufacturing. The position supports the business unit’s strategic goal of building a world-class, fully integrated solar manufacturing ecosystem by implementing advanced security measures and leveraging automation tools to enhance operational efficiency. Source: Adani Group | Job ID: 43993
About Anyscale At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role The Customer Engineer will play a crucial role in the customers’ post-sale journey - helping them to onboard, adopt and grow on Anyscale, troubleshooting and resolving open customer tickets and driving consumption. Anyscale is an ever evolving platform and hence will require close co-ordination with our engineering teams to debug complex issues. This is an exciting role for those who are technically curious and passionate about ML/AI, LLM, vLLM and the role of AI in next generation applications. It’s an opportunity to make a significant impact in a collaborative, fast-paced environment while building a new segment in this space. In this role, you’ll be able to Resolve customer issues and help in their successful adoption of Anyscale platform Be a technical advisor, and internal champion for our key customers Own customer issues end-to-end, from troubleshooting, triaging, escalations and eventual resolution Participate in our follow-the-sun customer support model to ensure continuity in resolving high priority tickets Keep track of open customer bugs and feature requests to influence prioritization and provide timely customer updates upon resolution Contribute towards improvement of internal tools an
Business Systems drives efficiency across Datadog through business process analysis, systems automation and integrations, AI agent and MCP development, and vendor/software review. The team is increasingly embedded in cross-functional initiatives across People, Finance, GTM , Legal, Recruiting, and Technical Solutions — translating ambiguous business problems into scoped, buildable solutions and owning delivery end-to-end. This is not a generalist BSA role. Each Senior BSA will own a cluster of business functions end-to-end, acting as an internal product manager for their domain rather than processing inbound requests reactively. There are two openings, each covering a different domain: People, Recruiting, Legal or Finance, GTM, Procurement As the role evolves alongside AI, Senior BSAs leverage agentic tools for data aggregation and context gathering while remaining the critical human-in-the-loop layer — owning business context and product outlook, validating use cases, managing stakeholder relationships, and making the judgment calls agents can't. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Discovery and scoping: Translate ambiguous asks from business stakeholders into well-defined requirements that Business Systems Engineers can build against. Many stakeholders don't know what they want or the cross-functional impact of what they're asking for — you surface both before engineering begins. Cross-functional visibility: Identify dependencies, downstream impacts, and integration considerations that requesting teams miss. Proactive opportunity identification: Develop deep domain knowledge to identify automation and AI opportunities before they become inbound requests, shifting the team from reactive in
At NVIDIA, we push the boundaries of computing innovation. Our ASIC Verification Engineers focus on developing the world’s top SoCs and GPUs. Joining us as a Senior ASIC Verification Engineer - GPU means working on modern technology powering consumer graphics and AI applications. This position is ideal for those passionate about technology and eager to impact computing’s future. What you'll be doing: As a key member of our ASIC Verification team, you will verify the design and implementation of the industry's leading GPUs. You will be responsible for verifying the ASIC build, architecture, golden models, and micro-architecture using advanced verification methodologies such as UVM or equivalent. Understand the design and implementation of your unit/cluster/chip, define the verification scope, develop the verification infrastructure, and verify the correctness of the design. Collaborate with architects, designers, and pre- and post-silicon verification teams to accomplish your task. What we need to see: Bachelor's Degree in EE, CS, or CE or equivalent experience. 5+ years of relevant experience. Experience in verification using random stimulus along with functional coverage and assertion-based verification methodologies. Experience with design and verification tools (VCS or equivalent simulation tools, debug tools like Verdi, Indago, GDB). Expertise in System Verilog or similar HVL. Strong debugging and analytical skills. Perl and C/C++ programming language experience desirable. Strong communication skills and the ability & desire to work as a great teammate are huge pluses. Experience in crafting test bench environments for unit and system level verification. #LI-Hybrid Your base salary will be det
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Business Strategy COVEC Manager Business Strategy COVEC Overview Reporting directly to the Vice President, Business Strategy, South LAC, with a dual reporting line to the COVEC Cluster Lead, this role leads Mastercard’s strategy agenda across Colombia, Ecuador, Venezuela, Guyana, and Suriname, while supporting broader South Division priorities. As a member of both the South LAC Strategy team and the COVEC Leadership Team, the role partners across the organization to shape strategic priorities, identify new growth opportunities, evaluate investments and business initiatives, support key decision-making processes, and drive execution against critical business objectives. The position works closely with leadership teams across markets, products, services, operations, finance, legal, public affairs, people, and other functions to ensure alignment, mobilize resources, and deliver measurable business impact. The role serves as a trusted advisor to senior leadership, helping translate market opportunities, industry trends, and business challenges into actionable strategies that accelerate growth, strengthen Mastercard’s competitive position, and support the delivery of cluster and divisional objectives. Key Responsibilities • Lead the development and execution of Mastercard’s strat
We build and operate the compute infrastructure our researchers run on, supporting large-scale processing of historical market data and model training on our own hardware across multiple data centers. Our environment includes bare-metal Linux, virtualization, storage, and GPU clusters, where performance, reliability, and predictable system behavior are critical. Our Infrastructure team covers monitoring and automation, distributed storage, hardware and OS provisioning, GPU clusters and workload scheduling, high-speed networking, L2/L3 Linux support, and security engineering. Engineers here own their tasks end to end, so there's room to go deeper in your area and pick up the parts you haven't touched yet. We’re looking for a Linux Infrastructure Engineer who can work hands-on with server and cluster environments, from deployment and configuration to performance tuning, troubleshooting, and ongoing improvement What You’ll Be Doing: Deploying, configuring, and maintaining Linux-based bare-metal servers across our data centers Building and operating clustered environments, including virtualization, storage, GPU compute, and database clusters Troubleshooting complex Linux, hardware, networking, and cluster-level issues Performance tuning for throughput, latency, stability, and resource utilization Monitoring infrastructure health and performance, identifying bottlenecks, and preventing recurring issues Supporting the full server lifecycle: provisioning, setup, upgrades, and maintenance Improving reliability and predictability during failures, maintenance, and scaling Automating provisioning, configuration, and operational tasks, primarily using Ansible and scripting What We Look For In You: Strong hands-on Linux administration and troubleshooting experience Production experience with on-premise, bare-metal infrastructure Good understanding of Linux performance and bottleneck analysis Experience with: infrastructure monitoring and troubleshooting production issues,
About the Role: We're hiring Senior and Staff Data Platform Engineers to join the Data Infrastructure teams in Toronto. Together these teams own the infrastructure that processes billions of events per day: Spark-on-Kubernetes, Flink and Kinesis pipelines, a multi-petabyte Delta Lake, a large-scale MemoryDB feature store, Databricks multi-environment operations, and the catalog and lifecycle systems that govern it. The team is small and senior. Each engineer owns major platform components: you design it, build it, and support it in production. This is a hybrid-role based out of our Toronto office. You must be willing to travel to our Toronto office two days/week. What You'll Do: Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring Your Background: 3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling Production AWS experience or equivalent: EKS, S3, Kinesis, and mu
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the Role Anyscale runs on a small, high-leverage IT function, and we're looking for an IT Engineer to own it. This is a hands-on engineering role, not a help desk or tier-1 support role. Roughly half the job is software: you will own and extend a set of internal automation and notification services wired into our directory and hardware systems. You should be comfortable owning a codebase and a GitHub repository, not only vendor admin consoles. The other half is the identity, endpoint, hardware, and vendor backbone that keeps a company of roughly 200 people, close to half of them engineers, productive across three public clouds. Reporting to the Head of Security and IT, you will work closely with security, engineering, and the people team. This role is based in the San Francisco Bay Area on a hybrid schedule. What You'll Do Own identity and access: Okta and Google Workspace administration, group-based authorization, 1Password Enterprise, passkeys, and automation driven provisioning and deprovisioning. Own and extend our internal automation services, and run them on scoped service accounts with sound security and hygiene. Own the endpoint fleet through mobile device management, including device trust posture checks, endpoint
About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for a Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud as easy as on your laptop. As part of the Infra team, we build the scalable, secure, and robust backbone that enables this vision. Our team is responsible for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which ensures high-performance execution of distributed workloads. We are seeking a talented engineers with a strong background in control plane and data plane development, along with expertise in Kubernetes, container orchestration, and cloud-native infrastructure. You will play a crucial role in designing, implementing, and optimizing the critical infrastructure that powers Anyscale’s cloud platform. You will have the opportunity to work on open-source Ray, contribute to our infinite laptop proprietary product, and develop seamless integration between the two, while also delivering high-impact features for our customers. A snapshot of projects you may work on Design, build, and scale services that orches
Get new cluster chef jobs by email
Daily job updates · Unsubscribe anytime