Jobs in United States

Fleet Operations Associate in United States

113 active opportunities · Updated October 2026

Explore current fleet operations associate jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

$342K – $445K/yr

Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec

TypeScriptPythonJavaSQL
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Client Platform Engineering (CPE) team delivers trusted devices at scale: secure by default, reliable by design, and effortless to use. We own platform capabilities across macOS, Windows, iOS, Android, and Linux, spanning endpoint posture and device trust, application delivery, onboarding, updates, telemetry, workflow orchestration, and employee-facing remediation. The team partners deeply with Security, Research, Applied, and specialized engineering groups to enable and protect OpenAI while reducing friction for the people advancing our mission. About the Role As an Engineering Manager for CPE, you will lead a team of engineers responsible for the strategy, delivery, and operation of OpenAI’s cross-platform client foundation. You will combine people leadership with strong technical judgment: setting direction, developing engineers, reviewing architecture and tradeoffs, and creating the operating mechanisms that turn ambiguous needs into durable platform outcomes. This is a high-leverage role at the intersection of security, reliability, developer velocity, and employee experience. CPE is a highly technical platform engineering organization delivering first-party services, automation, observability, and safe fleet operations. We’re looking for a leader who can guide its next chapter, scaling the team and its systems, partnering across the company, and raising the bar for secure, reliable, low-friction experiences across every supported platform. In this role, you will: Lead and develop a high-performing engineering team; hire thoughtfully, coach engineers, create clarity, and foster an inclusive, high-accountability culture that pushes perceived limits. Define and execute a multi-year client-platform strategy and roadmap across macOS, Windows, iOS, Android, and Linux, including how Codex and agents can reshape employee computing. Provide technical direction for endpoint posture, device trust, application delivery, device onboarding, updates,

AWSKubernetesCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching

AWSAzureRestAI
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -73.6%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments. We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment. This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.. RESPONSIBILITIES Core Responsibilities: Drive suppliers to keep the maximum amount of the GPU fleet online and healthy. Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs. Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier. SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short. Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability. Scope and

Machine LearningAIGoFinance
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. Lab Operations is the physical backbone of that work: we build and maintain the device fleets, test environments, and hardware-in-the-loop labs that let teams test repeatably, understand failures, and ship with confidence. About the Role As a Lab Operations Manager, Systems Integration , you will own the day-to-day operation of a large-scale consumer device test lab. This is a hands-on operations leadership role: you’ll keep device fleets, test rigs, lab infrastructure, inventory, provisioning, maintenance, and logistics running smoothly so engineers and QA technicians have reliable environments for validation and release testing. We’re looking for someone who is highly organized, technically hands-on, comfortable with consumer electronics and lab equipment, and experienced operating complex physical test environments at scale. Because this is a new category of devices, you’ll have the opportunity to build the lab operating model early—shaping the systems, standards, and workflows that support products from prototype through launch. In this role, you will: Own device fleet and inventory: Manage configuration, deployment, tracking, lifecycle, and accurate asset records for a large fleet of consumer devices and test

Artificial IntelligenceAILogistics
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.6%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are looking for an IT Support / Operations Engineer to join Baseten as we continue to scale our IT team. In this role, you will play a critical part in bringing our technical support entirely in-house to provide a seamless, high-touch experience for all Baseten employees. As we continue to scale, you will be the primary point of contact for day-to-day technical issues, allowing you to have a direct impact on our team's productivity and overall office environment. This position is ideal for a hands-on problem solver who enjoys a mix of hardware and software troubleshooting, user lifecycle management, and maintaining the physical IT infrastructure of a modern office. While you will focus heavily on elevating our internal support standards, you will also assist with systems administration and workflow automation as our company evolves. This is a hybrid role based out of our San Francisco or New York office, following our standard policy of three days per week in-person to ensure our physical office and AV systems remain high-performing and reliable. RESPONSIBILITIES Serve as the escalation point for day-to-day technical support, diagnosing and resolving hardware and software issues across our Mac and Windows fleet Manage user lifecycle administration including provisioning, deprovisioning, and access management across all systems and services Own the IT onboarding experience for new employees — from laptop set

Machine LearningAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer leading Fleet Management, you will be the overall technical lead across three pods and the person who sets the technical direction for the fleet management layer of Roblox. This is a hands-on, deeply technical leadership role that owns all of Roblox's compute capacity end to end: from low-level provisioning and the data plane, up through the control planes that operate it, and all the way to the UI and internal-facing products that let teams self-serve capacity. Your org centralizes security, maintenance operations, and the uptime of every Roblox Kubernetes cluster, and governs the internal customer contracts that drive automation across the fleet spanning Roblox data centers and cloud providers. You will guide architecture, raise the engineering bar, and make sure compute capacity supply and demand stay in balance as the fleet grows. You will: Serve as the overall technical lead for three Fleet Management pods, setting and aligning the technical direction across low-level provisioning, the data plane, and the control plane and product surfaces above them. Architect the declarative, Kubernetes-style control planes that operate Roblox's compute fleet across o

SQLAWSKubernetesGit
Y
📍 New York, NY, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Us: YipitData is the leading market research and analytics firm for the disruptive economy and recently raised up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points every day to provide accurate, detailed insights on ridesharing, e-commerce marketplaces, payments and more. Our on-demand insights team uses proprietary technology to identify, license, clean and analyze the data many of the world’s largest investment funds and corporations depend on. For three years and counting, we have been recognized as one of Inc’s Best Workplaces . We are a fast-growing technology company backed by The Carlyle Group and Norwest Venture Partners. Our offices are located in NYC, Austin, Miami, Denver, Mountain View, Seattle , Hong Kong, Shanghai, Beijing, Guangzhou, and Singapore. We cultivate a people-centric culture focused on mastery, ownership, and transparency. About the Role: This is a hybrid role based in our New York City headquarters. Employees are expected to work from the NYC office three days per week. We expect East Coast working hours. As Our IT Delivery Engineer, You Will: Design, implement, and maintain the systems and platforms that support the company’s internal IT environment Administer and troubleshoot SaaS platforms and endpoint management systems including Kandji, Google Workspace, Okta, Slack, Zoom, and other core business tools Manage and improve device management and endpoint configuration across the global fleet using MDM platforms Partner with Security and Infrastructure teams to ensure systems meet company security and compliance standards Lead and contribute to IT infrastructure and service delivery improvement projects that increase system reliability and operational efficiency Evaluate and implement new technologies that enhance IT service delivery, system management, and automation Develop automation and tooling to streamline provisioning, configuration management, and operational workflows Ma

O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models. We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience. You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments. Key Responsibilities Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency. Develop forecasting models for inference demand across products, regions, and model families. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs. Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions. Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps. Communicate technical findings clearly to both engineering teams and executive leadership. Qualifications MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience). 5+ years of experience working in the infrastructure data science space. Strong ex

PythonSQLAWSRest
L
📍 San Diego, California, Canada
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

The Undersea Systems Division at Leidos currently has an opening for a Technical Deputy Program Manager to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Technical Deputy Program Manager and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, a

LinuxProject ManagementPmp
L
📍 San Diego, California, Canada
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

The Undersea Systems Division at Leidos currently has an opening for a Senior Systems Engineer to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Senior Systems Engineer and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, action-item logs, a

LinuxProject ManagementPmp
L
📍 San Diego, California, Canada
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

The Undersea Systems Division at Leidos currently has an opening for a Systems Engineer to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Systems Engineer and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, action-item logs, and status rep

LinuxProject ManagementPmp
🔔

Get new fleet operations associate jobs in United States by email

Daily job updates · Unsubscribe anytime