About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil
Jobiba hiring network
Fleet Operations Associate Jobs
388 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current fleet operations associate jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Location : Come and join us either in Berlin or Hamburg! Freenow by Lyft empowers smarter mobility decisions helping people to move freely and cities to thrive. One of our core goals for 2026 is the complete digitalization of the European taxi market. We are partnering with local fleets to provide them with our advanced Dispatch and Fleet Management software completely for free, helping them cut costs, defend against ride-hailing giants, and thrive. We are looking for an empathetic, relentless, and hands-on Sales Manager to be the face of this initiative! YOUR DAILY ADVENTURES WILL INCLUDE: Prospect, engage, travel and pitch our Dispatch tool to SMB taxi companies. You will be the primary catalyst for transitioning legacy fleets into modern, connected partners Guide fleet owners and head dispatchers through the fear of operational change Act as the voice of our partners. You will gather on-the-ground feedback about our software and collaborate directly with our Product and Strategy teams to help "co-build" the ultimate solution TO BE SUCCESSFUL IN THIS ROLE: 3-5 years in B2B Business Development, Sales, Account Executive roles. Backgrounds in mobility, logistics, or selling SaaS to SMBs are highly preferred Proven track record of managing a B2B pipeline and handling objections. You understand that even when a product is free, you still have to expertly "sell" the time, effort, and operational shift required to adopt it Strong communication skills with the ability to build trust and long-term relationships Ability to deeply understand how a partner's business operates on a procedural level. You are adept at guiding new partners step-by-step through the transition, ensuring they smoothly adopt new software and daily routines You view travel not as a chore, but as an essential part of the job. You are comfortable being on the road for a significant portion of your time (up to 40-50%) to meet partners exactly where they operate Fluency in German and English is required
About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
At Jamf, we believe in an open, flexible culture based on respect and trust. Our track record and thriving work environment all stem from the freedom we grant ourselves to get the job done right. We take pride in helping tens of thousands of customers around the globe succeed with Apple. The secret to our success lies in our connectivity, while operating with a high degree of flexibility. Work-life balance remains our priority while feeling connected is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. This role is offered as a hybrid in Brno, Czech Republic . We are only able to accept applications for those based in the Czech Republic or who have sponsorship to live and work in the Czech Republic . What you'll do at Jamf : At Jamf, we empower people to be their best selves and do their best work. You will join Ocean, the team that owns "Blueprints" - Jamf's implementation of Apple Declarative Device Management (DDM). Blueprints is how Jamf administrators declare the desired state of their Apple fleet and have it delivered reliably to hundreds of thousands of devices. This is a primarily frontend role. You'll spend most of your time in our micro-frontend layer, building the interfaces admins use to compose and manage Blueprints. Backend development is not expected: over time you'll grow comfortable reading and debugging the Java services your UI depends on, but that's a bonus, not a day-one requirement. You'll work in a close-knit agile team, developing new features, breaking down customer problems, and owning the areas you touch. As part of an international ART, you'll collaborate daily with engineers in Czechia and Poland, and join a team that knows the Blueprints domain deeply and is invested in bringing new members up to speed - with a clear onboarding path, a manager who runs regular 1:1s and cares about your growth, and space to learn Apple device management on the job. What you
We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Bengaluru for our hybrid working model. Position Expectations Own the reliability architecture of the platform across regions and cloud providers Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices Set operational standards for the team: on-call quality, incident response, SLO discipline Mentor and technically develop the SRE team Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Qualifications 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms Proficiency in Python, Go, or a similar programming language Understand workload isolati
We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB. As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition. We are looking to speak to candidates who are based in Gurugram for our hybrid working model. Position Expectations Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing Identify and configure key metrics to detect incidents and quantify service health, availability, and performance Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure Mentor early-career SREs and contribute to the team’s operational practices as it grows Qualifications Strong background in software development and operating distributed systems 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues Expertise in cloud infrastructure platforms, in
About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructur
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with engineering teams of key CSP / hyperscale customers to ensure they can reliably manage, update, and operate NVIDIA GPU firmware at fleet scale. You will drive work streams with engineering teams of key CSPs/hyperscale customers to build shared understanding of GPU firmware and system software integration, incorporate their feedback into NVIDIA's feature roadmap and delivery plan, and ensure customer-side automation and recovery procedures are ready before each firmware release. Your cross-CSP visibility enables you to identify patterns in GPU firmware operational challenges that drive systemic improvements no single customer engagement could surface alone. What you'll be doing: Drive GPU firmware & siftware work streams with CSP engineering teams — ensuring they understand GPU firmware architecture (VBIOS, InfoROM, microcontroller firmware), update sequencing, recovery procedures, and GPU power management Gather and synthesize CSP feedback on GPU firmware/software — covering manageability, observability, security requirements (e.g., multi-tenancy isolation, secure boot, attestation), and performance — and champion those priorities into NVIDIA's GPU firmware/software feature roadmap and delivery plan Drive GPU firmware update orchestration for large-scale deployments — multi-GPU update sequencing, rollback strategy, failure handling, and validation across hundreds of GPUs per rack Serve as the technical focal point between NVIDIA and CSP firmware/software engineering — ensuring GPU behaviors (error recovery flows, thermal protection, power state transitions) are well-documented and accessible for customer integration Identify cross-CSP GPU SW/FW issue patterns — common update failu
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
As a Staff Engineer on Datadog's Compute – Disruption and Workload Placement team, you'll help define how our Kubernetes fleet scales to meet the demands of rapidly growing AI and cloud-native workloads. You'll work on the systems that ensure engineering teams have the right compute capacity, in the right region, at the right time across AWS, Google Cloud, and Azure. This is a highly technical, high-impact role where you'll shape the future of capacity orchestration, influence platform architecture, and solve infrastructure challenges that directly support Datadog's continued growth. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers. Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance. Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale. Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency. Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies. Who You Are: You have significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems. You are an experienced software engineer with strong programming skills, ideally in Go or a comparable systems programming language. You have hands-on experience with at least one major cloud provider (AWS, Google Cloud, or Azure) and understand distributed cloud infrastructure. Yo
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. Summary: Lead the Switchboard team, which owns the Tenant management product for Dedicated environments. This team is diverse in skills (Frontend, Backend and Reliability engineers) and owns the product end to end. The product enables self-service and workflow automations for customers and internal operators of our large scale Dedicated tenant fleet. Key Responsibilities: Lead the fullstack Switchboard team responsible for GitLab Dedicated’s customer-facing control panel, and set clear direction for how the team delivers and evolves that product. There is a strong coordination with the assigned Product Manager. Hire and develop a
About the Team The ChatGPT Model Flywheel team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. Team Focus Areas Model Experimentation: Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management. Model Deployment: Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling. Automate capacity management and incorporate platform-wide health monitors. Model Measurement: Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards. Improve end-to-end feedback loops for continual model improvement. Key Partnerships Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API. In this role, you will: Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks. Drive expansion and improvement of multi-tier model experiences. Support and scale self-serve experiment capabilities and automated guardrails. Lead model rollout automation, capacity management, and health monitoring. Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.). You might thrive in this role if you have: Proven experience leading engineering teams in complex, cross-functional environments. Demonstrated success shipping production systems at scale (ideally for AI or large backend services). Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling. Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders. Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus. Why Work With Us Tackle highly impactful technical challenges at the cutting edg
Get new fleet operations associate jobs by email
Daily job updates · Unsubscribe anytime