Jobiba hiring network

Fleet Operations Associate Jobs

388 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current fleet operations associate jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
12 days ago

NVIDIA DGX Cloud is an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As the Distinguished Engineer for Security Architecture, within our Security Engineering organization, you will set the security design bar for an AI factory of hundreds of thousands of GPUs, and then build against it alongside the teams. This is the founding architecture seat in a new organization. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and this is the architecture function inside it. You will define the security design standard for DGX Cloud, a bar that sits above the company floor, and hold it from inside the teams doing the building. Security here is fleet horizontal and stack vertical, so your scope runs from the hardware root of trust and the hardened baseline, through tenancy and GPU workload isolation, to the services and APIs built on top, across every DGX Cloud engineering organization. A small team of Principal Engineers will report to you and hold the bar at domain depth. This is still a hands-on seat, and you stay in the design with them. You will also serve as DGX Cloud's technical interface into NVIDIA's central security organization. There is no architecture review board here and no approval queue; the bar holds because the strongest security engineers in the room helped set it and helped ship it. What You Will Be Doing: Set the DGX Cloud Security Bar: Own the security design standard across DGX Cloud (tenancy, GPU workloads, identity, supply chain, and isolation) and make it concrete. Reference architectures, golden paths, and requirements engineers can actually build against, not a policy library. Hold the Bar by Building: Embed with engineering teams on real work: join the design, learn the code, help ship the thing rather than grade it afterward.

REMOTEkuberneteslinuxartificial intelligence
View job →

Job Title Sales, Strategic Accounts Director - Image Guided Therapy Systems (Southeast/Mid-Atlantic) Job Description The Director of IGT Strategic Accounts will work with 2 IGTS Districts and be responsible for order and revenue growth at a targeted set of named accounts. Your primary responsibilities will be to act strategically and grow Philips Marketshare by focusing on: executing fleet replacement plans of Philips Install Base accelerating competitive replacements at accounts where Philips is not installed (red accounts). In this role, you will work closely with the Philips IGT Sales and Marketing teams, you will build deep customer relationships at multiple touchpoints throughout your assigned health systems, all while developing and deploying strategies to solve interesting hospital challenges to help IGT continue to win. Your role Deliver agreed-upon growth in targeted accounts, and contribute to the overall sales results of the covered districts. Actively manage, monitor and continually improve the team’s overall sales process to ensure a successful productivity ramp in order to exceed revenue and commercial goals. Analyze and quantify equipment replacement potential at targeted accounts, assess regional market conditions, and use this to create and execute strategies to help grow Philips share Partner with CADs and ECEs to relentlessly assess and build out IDN and GPO sales strategies that help the district grow Philips IB or replace competitive IB Track and analyze customer performance to identify positive or negative trends, all while looking for opportunities to find mutually beneficial areas for growth. Able to quickly digest data, identify trends, and turn them into proactive strategies Ability to both work with and lead the IGT field sales teams in a matrix environ

SA
Scale AI
📍 San Francisco• Full-time• From $252K/yr
16 days ago

The Public Sector software engineers (SWEs) create the core product building blocks forward-deployed teams use to develop agentic capabilities that function across multiple domains. SWEs responsibilities include building the systems required to ingest and process federal datasets to support real-time decision-making in contested environments. We develop novel agentic enabling capabilities that includes: Create multi-layered guardrails around agents Optimize data retrieval for agents Orchestrate fleets of asynchronous agents Automatically alerts users to deviations in data Illustrating how an agent reached a decision As a Staff Software Engineer, you will orchestrate the implementation of vertical features and horizontal capabilities to include mentoring other engineers on defining requirements with stakeholders and communication tradeoffs of technical implementations on feature and capabilities until they are accepted by the stakeholders. You will: Orchestrate feature implementation across the Federal engineering team to ensure architectural consistency. Define technical strategy for agentic guardrails, explainability, and fleet orchestration. Ensure system reliability and performance across multiple security classifications and network types. Mentor engineers in the process of defining requirements with stakeholders and gathering acceptance. Communicate high-level technical trade-offs and implementation strategies to senior government stakeholders and Scale C-Suite members. Influence the long-term product strategy and technical roadmap for the Federal business unit. Consult on the architecture of AI-powered solutions for large-scale federal contracts. Ideally you will have: Full Stack Development: Proficiency in front-end, back-end development and infrastructure, including experience with modern web development frameworks, programming languages, and databases Cloud-Native Technologies: Familiarity with cloud platforms (e.g., AWS, Azure, GCP) and experience in

awsazuregcp
View job →
S
Supabase
📍 Remote• Full-time• Remote
1mo ago

About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role We are seeking a platform engineer to join our Compute Capacity team. This team owns the capacity plan that keeps compute supply ahead of demand across every region we operate in — the forecasting, buffer policy, and provisioning systems that make sure a project never runs into a wall it didn't know was there. You'll work on the systems that turn a capacity plan into provisioned reality: reservation acquisition, fleet reconciliation, and the automation that keeps what we've committed to in sync with what we're actually running. You'll help build the metrics and alerting that let capacity problems surface months out, on vendor lead time, rather than at the moment someone needs the room. You'll design, build, and operate systems that are both robust and highly automated — helping us hold the right buffer at the right cost, catch drift before it becomes a shortage, and give every team a single, trustworthy view of how much room we have across the millions of databases we manage. What You'll Be Responsible For Help build and maintain the capacity plan that keeps Supabase's compute supply ahead of demand across regions and instance families Support buffer policy by modeling headroom targets and their cost tradeoffs for review and sign-off Build and maintain automation that turns the capacity plan into provisioned reality — reservation acquisition and top-up, fleet reconciliation, drift detection between committed and running capacity Extend our infrastructure as code for capacity-relevant provisioning Instrument capacity: build and maintain metrics for saturation, reservation coverage, idle buffer, forecast error, and provisioning latency Build and tune capacity alerting so headroom,

REMOTEtypescriptpythonaws
View job →

As Senior Data Scientist for Engineering Systems you will work independently alongside sharp, generous, and pragmatic engineers from Server Query, Atlas Clusters, and Release Quality, among other teams. Together, we tackle problems spanning resource scaling across the Atlas fleet, safe feature rollout to MongoDB clusters, automated incident response and query engine performance. Join the Platform Data Science team and help us research, prototype and ship machine learning features for MongoDB’s core server, query engine and Atlas, our database-as-a-service cloud offering. We are looking to speak to candidates who are based in Dublin or Cork for our hybrid working model. What You'll Do Partner with Server Query, Atlas Clusters, Release Quality and other engineers to embed algorithmic rigor and optimization into resource scaling, release-safety and monitoring systems across the fleet and inside query engine Deliver production-ready, thoroughly tested statistical and ML algorithms with well-identified limitations that deliver measurable business impact, not just an impressive-sounding methodology Own the full feedback loop: instrument model architecture with the metrics needed to track performance and create dashboards in collaboration with our stellar analytics team, collect feedback from users and metrics to diagnose issues or opportunities, and iterate accordingly Deliver thoughtful, kind code reviews to your peers and act as a core contributor to internal packages, tooling, and team processes that increase developer productivity Measures of Success In 3 months, you’re familiar with our workflow, have an elementary understanding of our product and what teams we work with. You have delivered small-to-medium improvements to our project portfolio In 6 months, you’ve delivered one feature you researched and prototyped from scratch and demonstrated its impact on business metrics of your choice In 12 months, you've established a track record of shipping ML-driven improveme

pythonmongodbaws
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
1mo ago

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by providing

pythonjavagit
View job →
D
Datadog
📍 New York• Full-time• From $192K/yr
1mo ago

Senior Software Engineer - Streaming Platform Client Data streams are mission-critical at Datadog, powering near real-time communication across the vast majority of our services. Our Streaming Platform group builds the core infrastructure and abstractions that ensure Datadog remains a trusted partner for engineers worldwide. See our blog post . The Streaming Platform Client team sits at the heart of this ecosystem. We own the Rust client library (producers and consumers) with language bindings for Java, Go, and Python. We focus on building intuitive APIs and abstractions that make a powerful distributed system easy to adopt and operate for the hundreds of internal users of our library. Our library runs on critical data paths that handle hundreds of millions of messages per second making performance, observability, and reliability paramount. We also develop and operate the service that bridges the clients fleet with the platform's control plane, handling complex balancing, scaling, and static stability challenges. We are seeking a Senior Software Engineer to help us evolve these features. You will collaborate directly with our users, tackle performance-critical code, and solve complex distributed systems challenges across the control plane, client libraries, and data plane. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Work within a distributed, high-impact team spanning Europe and the US, building critical technologies that power data pipelines for dozens of internal teams and hundreds of services. Architect and implement resilient interactions between our client libraries and the control plane. Optimize our high-throughput, low-level streaming library to push the boundaries of performance and efficiency. Champion the developer experience by pro

pythonjavagit
View job →
A
Airbnb
📍 United States• Full-time• From $212K/yr
1mo ago

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Our web and API surfaces handle requests from guests and hosts alongside a growing volume of automated agents: AI assistants, crawlers, and scrapers. We build the systems that bring clarity to this traffic, combining in-house ML and vendor signals to decide in real time how to serve billions of daily requests. Anti-bot and anti-scraping detection is our most adversarial mandate, but the wider challenge is full traffic classification: building evaluation frameworks that tell legitimate automation apart from abusive actors, so high-stakes decisions hold up across the fleet. The Difference You Will Make: You will architect and maintain Airbnb’s end-to-end traffic classification ML systems, balancing high-performance model deployment with rigorous offline data pipelines. Success is measured by your ability to harden edge-traffic policies—targeting reduced bot-incident MTTM—and by establishing rigorous evaluation practices that ensure foundational signal accuracy and evasion-resistance across the fleet. A Typical Day: Own the complete lifecycle of traffic-scoring models, from problem framing to real-time deployment, managing the adversarial feedback loop to ensure high evasion-resistance and directly drive reductions in bot-incident MTTM. Architect robust offline-to-online pipelines that produce certified source-of-truth datasets, establishing rigorous evaluation frameworks—such as stratified benchmarks and leakage-prevention checks—to ensure every model improvement is empirically measurable and defensible. Execute model optimization within strict millisecond latency budgets at the

sqlgitmachine learning
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

Join the engineering teams that bring OpenAI’s ideas safely to the world!! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re building the observability product for OpenAI—from scalable infrastructure to a rich, AI-powered UI. Our systems ingest over petabytes of logs and billions of time series metrics across our fleet. We're now layering intelligence on top—think agents that summarize SEVs, auto-generate dashboards, or help engineers debug through notebook-like UIs. We’re hiring software engineers across the stack—infra, backend, and product. You’ll join a small, gritty team building both foundational infra and novel internal tools to make OpenAI's production systems reliable, performant, and observable. What You’ll Do Own core observability infrastructure, including distributed logging, time series, and trace storage Build AI-native tools that help engineers detect, understand, and resolve issues autonomously. Contribute to UI experiences like dashboards, notebooking, or interactive debugging Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product You Might Be a Fit If You: Have operated large-scale distributed systems in production. ( especially logging systems or some other time series databases) Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems. Have full-stack chops or product sensibilities—you're excited to build real tools people use. Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc). Bonus : built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc). Why This Team We’re b

awskubernetesrest
View job →
O
1mo ago

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

awskubernetesrest
View job →
N
10 days ago

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are a world - class autonomous driving hardware and software development team. We have the best platform and have maintained a leading position in the field of artificial intelligence. Next, we will continue to deepen our efforts in the autonomous driving field and strive to bring sustained growth to our customers. We're looking for a Senior Software Triage Engineer with strong technical ability to deeply understand architectures and strong scripting experience to automate and Triage methodology, and the leadership to encourage our engineering team. As a key member of our automotive group, you'll be working on the real time challenges outstanding to the automotive industry and our automotive products. This is a key role to support AV software agile iteration & development for the release of clients, together working with the engineering teams, understanding the AV stack deeply and providing data-based judgement and delivering high-quality AV software to end customer. What you’ll be doing: Work with Product, Engineering, Model/SW Devs, In-car testing, Fleet teams on test request and triage planning, do live triage with accurate analysis and debugging steps, present top issues by end of day. Deep understand

artificial intelligenceai
View job →
M
Modal
📍 San Francisco• Full-time
11 days ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll automate the integration of new capacity from a growing set of hardware providers; from auditing and benchmarking hosts and clusters, to maintaining our machine images, configuring GPUs, RDMA, networking, and storage, and getting machines into production. You'll build the automation that keeps the fleet healthy without human intervention: detecting bad GPUs, thermals, and disks. You'll dig into whatever is between the hardware and the software that runs on

pythonlinuxai
View job →
M
Modal
📍 San Francisco• Full-time
11 days ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll make cold starts feel local when the data is hundreds of milliseconds away, designing the caching, preloading, and peer-to-peer layers that hide object-store latency and keep public ingress off saturated uplinks. You'll own durability and cost at petabyte scale, from streaming and batch replication between origins, to garbage collecti

M
11 days ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll own the full lifecycle of a machine, from accepting and benchmarking new hardware from a growing set of providers, to network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation of unhealthy hosts. You'll manage a team of 3–8 engineers while staying hands-on across the stack which involves BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services, and you'll shape our long-

M
Modal
📍 San Francisco• Full-time
11 days ago

About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll set technical direction for the primitives that other teams (filesystems, training, sandboxes) build on, balancing durability, latency, throughput, and cost. You'll own the roadmap from today's hardest problems (garbage collection at petabyte scale, active-active replication, rate limiting that protects the upstream without wasting ut

🔔

Get new fleet operations associate jobs by email

Daily job updates · Unsubscribe anytime