About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
Jobs in United States
Computer Operator in United States
518 active opportunities · Updated October 2026
Showing
15 jobs
Explore current computer operator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui
About the Team Industrial Compute is building the infrastructure ecosystem that enables OpenAI to train and deploy increasingly capable AI systems at unprecedented scale. The organization operates across compute supply, demand, infrastructure, partnerships, and the physical and commercial systems required to make large-scale compute available. Industrial Compute Strategy & Operations serves as the connective operating layer across this ecosystem. The team works directly with senior leadership across Scaling, Finance, Partnerships, Research, and Infrastructure to translate ambiguous, high-impact challenges into clear strategies, scalable operating mechanisms, and decisive execution. This team is responsible for ensuring that OpenAI’s compute strategy evolves into durable competitive advantage by identifying systemic constraints, aligning stakeholders around critical decisions, and driving the operating mechanisms required to execute at scale. About the Role We are seeking a highly experienced Strategic Operations leader to help shape and operationalize OpenAI’s compute strategy across supply, demand, infrastructure, partnerships, and commercial strategy. This is a senior individual contributor role operating at the intersection of strategy, operations, infrastructure, and executive decision-making. You will work closely with compute leadership to identify the most consequential problems facing the organization, develop structured approaches to solving them, align stakeholders across the company, and drive initiatives from ambiguous concepts through execution. The role will span both strategic and operational work. You may develop long-range compute strategies and investment frameworks, evaluate build-versus-buy decisions, shape major commercial transactions, establish organizational planning mechanisms, or take ownership of a cross-functional initiative that does not have a clear organizational home. Success in this role requires exceptional judgment, analytical
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Infrastructure team is central to this mission, setting the core strategy and implementing the vision. From site selection to deployment to operations, this team sits at the intersection of commercial, technical, and operational domains, interacting with experts and executives inside and outside of OpenAI. We design and operate mission-critical facilities that support cutting-edge AI workloads at scale. About the Role We are seeking a Facilities Operations Lead to support the commissioning, deployment, and long-term operation of our next-generation AI data centers. This role bridges the interface between data center construction and hardware landing, ensuring seamless integration of mission-critical infrastructure with hardware deployment timelines. You will define and execute commissioning plans, support infrastructure bring-up, and take ownership of operations and maintenance for cutting-edge, large-scale, AI data centers. You will collaborate closely with design, construction, and hardware teams to define repeatable processes for new data center builds and lead hands-on operations to uphold the performance and reliability of our deployed infrastructure. Key Responsibilities Define and execute sequences of operations, commissioning steps, and bring-up processes for mission-critical data center facilities. Interface with the design and hardware teams to define deployment procedures tailored to each data center and hardware configuration. Oversee installation, commissioning, and operational readiness of large-scale data center campuses. Manage monitoring, maintenance, and quality control of the data center infrastructure, including high-performance liquid cooling systems. Develop on-site operations staffing strategy. Develop and enforce procedures for planed and unplanned downtime and SLAs for critical
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our Tensix Team is building the future of AI compute with a ground-up architecture centered on scalable RISC-V processors. As we push performance boundaries, we’re reimagining the frontend of our RISC-V cores to deliver major gains in programmability, efficiency, and developer experience. This is a rare opportunity to shape the CPU architecture at the heart of our AI platform and lead one of the most strategic technical efforts at Tenstorrent. This role is hybrid, based out of Toronto, ON, Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Experienced Microarchitect: 10+ years of deep expertise in CPU performance modeling and microarchitecture design. AI Workload Expert: Deeply familiar with the computational and memory bottlenecks of modern AI workloads, particularly Large Language Models (LLMs). Hardware-Software Co-Designer: Driven to architect custom instruction set extensions and validate their performance gains against real-world workloads. Ways to stand-out: Familiarity with open-source RISC-V cores, AI-based agentic workflow experience What We Need Profile & Analyze: Dissect cutting-edge AI workloads to identi
About the Team OpenAI’s Compute Strategy team is responsible for securing and scaling the core resources that power our research and products. We partner across engineering, finance, legal, and operations to identify, negotiate, and execute strategic partnerships that expand OpenAI’s capacity for compute, power, and data center infrastructure. Our mandate spans energy procurement, real estate development, colocation, cloud service providers, silicon and strategic supply chain, and infrastructure financing—ensuring OpenAI can grow with speed, resilience, and cost-efficiency. About the Role We are hiring several Business Development Lead, Compute Strategy positions focused on compute infrastructure. Each hire will bring deep expertise in one or more focus areas while collaborating across the broader infrastructure stack. In this role, you will source opportunities, structure partnerships, and negotiate high-value agreements across OpenAI’s infrastructure ecosystem. You will work directly with external partners and suppliers while collaborating internally with engineering, legal, finance, and operations to ensure we have the resources needed to support state-of-the-art AI systems. This role requires technical fluency, commercial judgment, and disciplined execution. Your work will directly shape how quickly, reliably, and efficiently OpenAI can bring new compute capacity online. Each hire will focus on building partnerships and executing deals in one or more of the following areas: Energy and Power: securing scalable and sustainable energy supply. Land and Real Estate: identifying and securing strategic sites. Colocation : evaluating and contracting for third-party data center capacity. Cloud Service Providers (CSPs): structuring partnerships with hyperscalers and specialized AI cloud providers. Silicon: building semiconductor partnerships to secure advanced silicon and resilient long-term supply. Fiber & Equipment: securing fiber & critical data center equipmen
About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati
About the Team OpenAI Finance is responsible for ensuring the organization is set up for success in pursuit of its mission. The Technical Accounting team plays a crucial role in helping OpenAI navigate complex, judgmental, and rapidly evolving accounting matters with rigor and clarity. We aim to bring both technical excellence and strong business partnership to some of the most novel accounting questions in the industry. About the Role As Senior Manager, Technical Accounting, Compute Infrastructure, you will lead the evaluation, documentation, and operationalization of complex accounting matters related to OpenAI's compute infrastructure, strategic investments, and other non-routine business activities.. This role sits at the intersection of U.S. GAAP technical accounting, infrastructure strategy, financial reporting, controls, and cross-functional execution. Key areas may include cloud compute arrangements, data center and colocation arrangements, lease accounting under ASC 842, power purchase agreements, strategic investments, consolidation evaluations under ASC 810, financial instruments, and other emerging or non-standard arrangements. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead technical accounting analysis for complex, judgmental, and non-routine transactions under U.S. GAAP. Evaluate accounting implications for compute infrastructure arrangements, including cloud compute, data center, colocation, lease, PPA, infrastructure procurement, and related commercial arrangements. Partner with Controllership, Tax, Legal, FP&A, Procurement, Infrastructure, and other cross-functional teams to assess the accounting implications of new products, commercial arrangements, strategic transactions, and business initiatives. Prepare and review technical accounting memoranda, position papers, and other auditor-ready documentation. Translate
At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an
£215K – £260K/yr
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role Cohere is seeking a Global Public Policy Manager to lead policy engagement on compute infrastructure, export controls, AI competitiveness, and sovereign AI strategies. Governments increasingly view AI as critical national infrastructure and are investing heavily in compute capacity, energy resources, and domestic AI ecosystems. This role will help position Cohere as a trusted partner in emerging discussions around AI infrastructure, national competitiveness, and sovereign AI deployment. Key Responsibilities Monitor and analyze developments related to AI infrastructure, data centers, energy policy, semiconductor policy, export controls, and national AI strategies. Develop policy positions on sovereign AI, compute access, digital sovereignty, and AI competitiveness. Support engagement with governments developing AI infrastructure investment programs and national AI initiatives. Collaborate with commercial, product, and corporate development teams on strategic opportunities involving public-private partnerships. Represent Cohere in policy discussions related to AI infrastructure, energy requirements, and technology c
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, o
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Compute team, you will be the technical anchor for Roblox's GPU and AI accelerator capabilities. This is a battle-tested GPU expert role focused on the machine management layer and above: how GPU hosts are made production-ready, kept healthy, and turned into reliable compute for the workloads that depend on them. You will own the hard problems that show up only at scale, from driver and firmware management to GPU health, reliability, and performance across a rapidly growing fleet of accelerators spanning Roblox data centers and cloud environments. You will set the technical direction for GPU compute and up-level the entire organization's GPU expertise. You will: Serve as the GPU technical leader for the Compute team, partnering across Kubernetes, Machine Bootstrap, Networking, and Cloud to drive GPU strategy end to end. Own the GPU host lifecycle above raw fleet management: driver, firmware, and CUDA stack management, GPU health and telemetry, and remediation of GPU-specific failures (XID errors, ECC, thermal, NVLink and fabric faults). Architect how GPU capacity is exposed to compute platforms, including scheduling, isolation, and integration with Ku
From $277.4K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Technical Program Manager for Compute Infrastructure, you will lead technical programs to develop and rollout next generation technology solutions to process a diverse spectrum of infrastructure workloads across Roblox services to enable users around the world to efficiently, reliably, and securely enjoy the Roblox Platform. You Have: Been a technical leader with domain expertise in systems and software used to process and manage Cloud and on-prem Infrastructure 7+ years of experience in software industry driving the build of large-scale infrastructure Experienced with establishing work relationships across multi-disciplinary teams and earning trust as a technical leader with all partners Experienced with identifying critical technical problems and opportunities Experienced with delivering end to end technical programs through roadmapping and reliable execution Able to turn specific solutions into systems that benefit the larger community in the long run Knowledgeable about user needs, scoping, planning, execution and delivery Flexible around process, using process as a tool when it makes sense for a team Experience with Kubernetes, Cloud platforms (e.g. AWS), and AI/ML infrastru
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer leading Fleet Management, you will be the overall technical lead across three pods and the person who sets the technical direction for the fleet management layer of Roblox. This is a hands-on, deeply technical leadership role that owns all of Roblox's compute capacity end to end: from low-level provisioning and the data plane, up through the control planes that operate it, and all the way to the UI and internal-facing products that let teams self-serve capacity. Your org centralizes security, maintenance operations, and the uptime of every Roblox Kubernetes cluster, and governs the internal customer contracts that drive automation across the fleet spanning Roblox data centers and cloud providers. You will guide architecture, raise the engineering bar, and make sure compute capacity supply and demand stay in balance as the fleet grows. You will: Serve as the overall technical lead for three Fleet Management pods, setting and aligning the technical direction across low-level provisioning, the data plane, and the control plane and product surfaces above them. Architect the declarative, Kubernetes-style control planes that operate Roblox's compute fleet across o
Other cities to consider
More places hiring for this role
Get new computer operator jobs in United States by email
Daily job updates · Unsubscribe anytime