ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re looking for a Recruiting Coordinator to help create a seamless, welcoming, and well-organized interview experience for every candidate who engages with our team. You’ll work closely with our recruiters to coordinate both virtual and in-person interviews, support executive involvement when needed, and ensure candidates have everything they need during throughout their interview process. This role is ideal for someone who thrives on operational excellence, loves solving logistics problems on the fly, and brings both warmth and precision to every interaction. RESPONSIBILITIES Work closely with recruiters and hiring managers to coordinate interview loops and debriefs for candidates and the internal team members conducting interviews Ensure every candidate has a smooth, well-communicated, and positive experience Manage logistics for onsite interviews, including candidate arrival and workspace setup Proactively identify and solve day-of issues, including last-minute changes or scheduling conflicts Communicate clearly and promptly with candidates and internal teams about interview logistics and updates REQUIREMENTS 1+ year of recruiting or HR experience Detail-oriented and operationally strong—you know how to keep things moving Clear and professional written and verbal communication skills Personable and warm—you're great at making candidates feel welcome and supported Ability to think on your feet and respond to
Jobs in United States
Fleet Coordinator in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current fleet coordinator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and
About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT infrastructure team is responsible for ensuring that our products can serve rapidly growing demand with the performance, reliability, and quality our users expect. This work sits at the intersection of product demand, model deployment, inference, research, fleet, and capacity. The team translates changing product and model needs into clear capacity decisions and safe, scalable launches. About the Role We are seeking a Technical Program Manager to lead the operating system for Chat capacity and model deployment. You will connect demand forecasting and capacity allocation with model readiness, rollout planning, launch coordination, and post-deployment learning. You will also own mode deployment beyond capacity by working with cross functional teams across research, post-training, inference and product to own mainline model deployment. You will bring structure to constrained-capacity decisions, improve the tooling and mechanisms teams use to prioritize demand, and help new models reach users safely and efficiently. Success requires technical depth, sound judgment under ambiguity, and crisp execution across product, research, infrastructure, and operations teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own cross-functional programs for Chat capacity forecasting, allocation, headroom planning, and constrained-capacity operations. Build durable intake, prioritization, and decision mechanisms that connect product demand and model requirements to available serving capacity. Partner
What we’re doing isn’t easy, but nothing worth doing ever is. Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. We’re hiring a Manufacturing Reliability Engineer to own production test for our robots at our contract manufacturer: you’ll design and run robust end-to-end test protocols, provision fleets of robots for production, and own the KPIs that define production quality. This role is based in Austin, TX. However, the position will require 50% travel to the Milwaukee, WI area and requires close collaboration across software, hardware, operations, and product engineering teams. Key Responsibilities End-to-end test process ownership. Create, validate, and maintain production test protocols and gating criteria from incoming inspection through final test and shipment. Provisioning of bots. Design and operate provisioning flows (imaging, firmware deployment, configuration, validation) and the tooling/fixtures needed to provision and handoff robots for production. KPIs and continuous improvement. Own key production metrics — First Pass Yield (FPY), cycle time, and test coverage — and drive continuous improvements to meet throughput and quality targets. Test automation & infrastructure. Architect, implement, and maintain automated test frameworks, harnesses, and test rigs used at the CM site. Ensure tests are stable, fast, and provide actionable failure data. Cross-functional escalation & RCA. Lead root-cause analysis for field and production failures; coordinate corrective actions with design, firmware, and CM engineering to close quality loops. On-site production leadership. Be the onsite technical authority at the contract manufacturer: train operators, debug failures on the line, and continuously refine processes with CM partners. What Success Looks Like Improved FPY and reduced rework rates across production builds. Reduced per
NA Fleet Management (Nashville, TN) — Nashville, Tennessee, United States. Apply via Workday.
Datadog is looking for a Senior Product Manager to help lead the evolution of our fleet and lifecycle management capability, the product surface that gives customers visibility into, and control over, the observability software running across their infrastructure. This capability manages the deployment lifecycle for core observability agents and OpenTelemetry collectors running on customer hosts and containers. The Senior PM will expand the scope of fleet capability to additional Datadog software components, making it the single place customers go to see everything running in their environment, at any version, in any deployment model, and to manage it remotely and safely at scale, for both human operators and, increasingly, AI agents acting on their behalf. This is a high-visibility, cross-functional role. You'll partner with multiple engineering teams and be responsible for defining and delivering a coherent, unified fleet experience across UI, API, and MCP for customers. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Own and evolve the product vision and roadmap for a unified fleet and lifecycle management capability spanning multiple product lines and deployment models. Define what "managed" means for each new software component as it's brought into fleet, balancing consistency of experience with the realities of each component's operational model. Drive a phased expansion plan, sequencing new components into fleet based on customer value, technical complexity, and dependency readiness. Partner closely with engineering leads across several teams to align on shared architecture principles to support disparate software components. Represent the voice of the customer for a capability that must work equally well for human operators using a UI and for AI
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades Supporting research workflows with service frameworks and deployment systems Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. Interface with researchers and product teams to understand workload requirements Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: Have experience with hyperscale compute systems Possess strong programming skills Have experience working in public clouds (especially Azure) Have experience working in Kubernetes Execution focused mentality paired with a rigorous focus on user requirements As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI resea
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer leading Fleet Management, you will be the overall technical lead across three pods and the person who sets the technical direction for the fleet management layer of Roblox. This is a hands-on, deeply technical leadership role that owns all of Roblox's compute capacity end to end: from low-level provisioning and the data plane, up through the control planes that operate it, and all the way to the UI and internal-facing products that let teams self-serve capacity. Your org centralizes security, maintenance operations, and the uptime of every Roblox Kubernetes cluster, and governs the internal customer contracts that drive automation across the fleet spanning Roblox data centers and cloud providers. You will guide architecture, raise the engineering bar, and make sure compute capacity supply and demand stay in balance as the fleet grows. You will: Serve as the overall technical lead for three Fleet Management pods, setting and aligning the technical direction across low-level provisioning, the data plane, and the control plane and product surfaces above them. Architect the declarative, Kubernetes-style control planes that operate Roblox's compute fleet across o
About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
Other cities to consider
More places hiring for this role
Get new fleet coordinator jobs in United States by email
Daily job updates · Unsubscribe anytime