About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
Jobs in United States
Computer Operator in United States
518 active opportunities · Updated October 2026
Showing
15 jobs
Explore current computer operator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking an experienced Principal Hardware Diagnostics Engineer to design and develop diagnostics software used to monitor hardware health and diagnose system-level issues across Graphcore’s AI infrastructure platforms. This role focuses on building diagnostics agents, tools, and analytics frameworks that enable engineers and automation systems to identify, isolate, and resolve hardware issues across blade-level servers and rack-scale clusters. The Team Graphcore is a globally recognised leader in Artificial Intelligence computing systems. The company designs advanced semiconductors and data centre hardware that provide the specialised processing power needed to drive AI innovation, while delivering the efficiency required to support its broader adoption. The Systems Engineering and Platform Validation team ensures Graphcore’s AI compute platforms are reliable, diagnosable, and operationally robust at scale. The team co
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
Technical Program Manager – Applied Infrastructure About the Team The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety. The Applied Infrastructure TPM team partners across engineering to lead foundational programs that ensure OpenAI’s infrastructure can meet current and future demand. About the Role We’re looking for a seasoned Technical Program Manager to drive critical infrastructure programs across the Applied organization. This TPM will focus on cross-cutting initiatives such as general compute capacity planning, process transformation, cost and quota attribution and optimization, and coordination across infrastructure and product stakeholders. There will also be focus on evolving OpenAI’s infrastructure to support growth, scale and new products. This work is core to how OpenAI manages and grows its infrastructure footprint in a disciplined, scalable way. Location: San Francisco, CA (Hybrid – 3 days/week in-office) In this role, you will: Serve as the DRI for complex infrastructure programs spanning CPU planning, orchestration, and other resource management domains (e.g. networking, storage). Build and operationalize systems to capture demand signals, model future capacity needs, and align infrastructure planning across internal teams and partners external to the company. Partner closely with Infrastructure, Product and Finance teams to forecast infrastructure usage patterns and ensure supply/demand alignment. Lead cost attribution and quota enforcement programs to promote stability and ensure equitable access to resources across teams. Drive simplification and standardization of infrastructure tooling and processes across Applied and Infra organizations. Drive cross functional programs to evolve our infrastructure to support new growth and scale Work with external v
High Performance Workstation Business Development Manager- AMD Description - Sales and Technical Consultant, High-Performance Workstation Segment We are seeking a customer-facing AI and Acquisition subject matter expert role on the HPI Advanced Compute Solutions (ACS) sales team and will focus on growing HP’s Z high-performance workstation Total Addressable Market (TAM) within the AI PC market. This individual will be responsible for working with customers and partners in support of attracting new customers and growth in HP’s high performance compute solutions . The role analyzes market competition, gathers customer feedback, and shares insights with internal teams for improvement. The ideal candidate will possess a breadth of abilities and skills, including client hardware knowledge, AI and machine learning applications and solutions, strong marketing and presentation skills, solid understanding of customers’ business and decision makers, and strong team leadership to drive growth with an emphasis on Advanced Micro Devices (AMD) platforms with customers and partners. Most importantly, the role offers the opportunity to be a business leader in support of the sales teams nationally - contributing to the strategy, setting direction, and achieving success in AMD and HP. The role requires someone to be self-driven, capable of operating autonomously through ambiguity, and constantly striving for excellence. THE PERSON: We seek an individual with exceptional technical & business knowledge and familiarity with AI trends in the PC market. This individual should be comfortable working: Dynamic environment and have experience with PC client and AI hardware, software and systems. Work independently within a core team is important, and you will rely upon your excellent communication and relat
We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver communication libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. DL and HPC applications have a huge compute demand already and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! We are looking for a technical leader to manage our NVSHMEM and UCX libraries. This is an outstanding opportunity to push the limits on the state-of-the-art and deliver platforms the world has never seen before. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Lead, mentor, and grow your library engineering team and be responsible for the planning and execution of projects as well as the quality, and performance of your libraries. This is a technical leadership role so you will participate in feature design and implementation. Interact with internal and external partners and researchers to understand their use cases and requirements. Collaborate with engineering teams, program and product management, and partners to define the product roadmap. Continuously review and identify improvement opportunities in established processes, infrastructure, and practices to ensure the teams are executing in the most efficient and transparent manner. What we need to see: 10+ overall years of experience in the software industry with specialization in HPC networking or system software. 4+ years of management experience. BS, MS, or Ph.D. in C
NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect -- for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Industrial Compute team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role The Community Engagement Lead will be the primary bridge between OpenAI and the communities in Ohio. This role ensures that OpenAI builds strong, trust-based relationships with local stakeholders, communicates proactively about our projects, and integrates community priorities into our development approach. The role spans engagement, communications, and reputation management, and will partner closely with the Economic Development and Environmental leads. Key Responsibilities Build and maintain relationships with local leaders, community organizations, NGOs, and residents. Develop and execute community engagement strategies for new and existing sites. Represent OpenAI in public forums, hearings, and community events. Partner with the Economic Development Lead on incentive compliance and community benefits. Partner with the Environmental Lead on communicating environmental stewardship and sustainability efforts. Develop proactive communications to address concerns, highlight benefits, and reduce risk of opposition. Monitor community sentiment and advise executives on risks and opportunities. Create a community engagement playbook that can scale across geographies. Qualifications 8+ years in community affairs, public engagement, or corporate communications. Proven track record engaging diverse community stakeholders for large infrastructure or technology projects. Strong public speaking and facilitation skills. Ability to manage sensitive political and reputational issues. Experienc
About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling novel and consequential legal issues in AI. The commercial infrastructure legal team supports the systems and facilities that make advanced AI possible, including compute, first-party silicon, colocation, and first-party data center development. About the Role We’re seeking a senior commercial infrastructure lawyer to lead legal work across our rapidly growing data center development and colocation portfolio. You will advise on complex, high-value transactions involving new data center sites, power, construction, colocation, and related infrastructure arrangements. This role is designed for a lawyer who understands how data centers are developed and operated—not a general real estate practitioner. You will work closely with infrastructure, finance, procurement, and operations teams, while coordinating outside counsel where appropriate, to help projects move quickly and responsibly. San Francisco is preferred, and occasional travel may be required. In this role, you will: Lead commercial legal strategy and risk management for data center development and colocation transactions. Draft, negotiate, and advise on colocation agreements, construction agreements, power-related arrangements, and other contracts supporting data center development. Support first-party projects involving raw land, site and power rights, design and construction, EPC models, third-party build-to-suit structures, and leasebacks. Partner cross-functionally with infrastructure, finance, procurement, operations, and other stakeholders on fast-moving, complex projects. Manage and collaborate effectively with outside counsel across a high volume of infrastructure matters. Develop scalable legal playbooks and contracting approaches for a growing portfolio of data center sites. You might thrive in this role if you have: 12+ years of legal experience, primarily in commercial legal roles. Deep, hands-on experience
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Industrial Compute team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role OpenAI is seeking a Real Estate Lead to own land acquisition strategy and site-control execution for our next-generation data center portfolio across the U.S. This role sits at the front end of infrastructure delivery, translating market intelligence, broker/developer relationships, and commercial judgment into credible shovel-ready positions that meet OpenAI’s power, land, permitting, and expansion requirements. The Real Estate Lead will not operate as a traditional transactional real estate function. The scope is to help shape where OpenAI can build, secure the right land positions early, structure defensible commercial terms, and coordinate the legal, technical, and development work required to move opportunities from sourced lead to controlled site and ultimately to readiness handoff. This is an individual contributor lead role and does not have direct reports initially. The role also owns market prioritization and portfolio-level acquisition strategy across target geographies, and is expected to run multiple negotiations in parallel while translating site-control work into executive-ready acquisition recommendations. Key Responsibilities Own market prioritization and portfolio acquisition strategy across target geographies, including scenario analysis for speed, scale, expansion potential, and risk-adjusted economics. Proactively source and evaluate land parcels at scale to support long-term data center growth across priority markets. Build and manage a national pipeline of
About the Team OpenAI’s Strategic Finance organization provides financial insights and guidance to support the company’s long-term strategy and ambitious growth. We work across Partnerships, GTM, Product, Compute, and Operations to allocate and deploy resources to the highest-impact opportunities while protecting sustainable unit economics. The B2B Strategic Finance team focuses on the financial performance of our B2B products and GTM functions, ensuring tight alignment between financial objectives and company strategy. We: Drive operational planning, financial forecasting, and performance management for our B2B business, Provide analytically-grounded insights on B2B product and financial performance to inform strategic resource allocation, Build the “0→1” financial foundations required to scale and accelerate growth. Within B2B, Partnerships are a critical lever to accelerate growth. You will help shape this business and drive successful outcomes from it. About the Role This Strategic Finance Partnerships hire will own our B2B Partnerships P&L and help drive strategic decision-making. You will be a critical finance partner to our Partnership teams, deal leads, Product owners, and Operations teams, responsible for pricing, structuring and driving success from partnerships, translating opportunities into decision- and exec-ready economics, and managing the overall Partnerships book of business. You will shape how OpenAI approaches Partnerships revenue targets, unit economics, investments, measurement, and prioritization. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Help manage the overall Partnerships business at a strategic priority, top-line impact, and resourcing / investment level. Own the financial evaluation of new B2B partnerships, modeling revenue, contribution margin, compute costs, cash flow, investment needs, and risk. Driv
About the Team OpenAI, in close collaboration with our capital partners, is embarking on a journey to build the world’s most advanced AI infrastructure ecosystem. The Industrial Compute team is central to this mission, setting the core infra strategy and implementing this vision. From site selection to the buildout process, this team sits at the intersection of commercial, technical, strategy, and operations, interacting with teams and executives inside and outside of OpenAI. About the Role Responsible for validating that proposed sites are buildable, compliant, and cost-effective. You will lead diligence across civil, geotechnical, environmental, and entitlement dimensions, identifying risks and driving mitigation strategies. Key Responsibilities Lead all technical diligence: geotech, soils, title/ALTA surveys, mineral rights, and access. Oversee permitting/entitlement path and schedule governance with agencies. Evaluate generator air permits, wetlands, floodplain, and stormwater constraints. Manage consultants performing feasibility studies and environmental assessments. Deliver go/no-go recommendations with risk and mitigation options. Build diligence templates and playbooks to scale future site reviews. Qualifications 8+ years in land development, civil/environmental engineering, or data center diligence. Knowledge of permitting, entitlements, and AHJ engagement. Strong project management and technical review skills. Experience managing consultants and interpreting complex studies. Regularly communicate site readiness updates, risks, and milestones to executive stakeholders Establish and track key performance indicators to assess the effectiveness of the site selection program and the contributions of external vendors and partners. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely depl
Other cities to consider
More places hiring for this role
Get new computer operator jobs in United States by email
Daily job updates · Unsubscribe anytime