About the Team The Strategic Initiatives & Operations team is in need of a Technical Program Manager (TPM) to streamline our processes, including full safety governance and integration of various safety research and mitigations into our ChatGPT, API, and any frontier models. This role is critical for driving safe deployment of our new models, synthesizing inputs from multiple stakeholders, ranging across research, product, engineering, legal and policy, and ensuring all the risks are effectively and properly monitored, mitigated or resolved. About the Role As a TPM, you will be responsible for critical tasks ranging from tracking safety research progress and risk tables to overseeing the quality of human data campaigns – acting as the connective tissue to enhance the deployment of OpenAI’s safety system. Additionally, you will create and execute a compute roadmap for your team to ensure that our top priorities are resourced while taking advantage of new opportunities to make key safety research discoveries. Your primary focus will be to ensure our models are qualified for safe deployment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Manage key risk areas and corresponding stakeholders. Keep track of a stack of existing and future mitigations for every major product and model deployment. Standardize the lifecycle of risk assessment, setting safety bars, consolidating inputs from multiple stakeholders across research, product, engineering, legal and policy, pre-launch safety reviews and post-launch followup. Manage pre-launch safety reviews. Share launch calendars and key safety practices and evaluations with our key parter (i.e. Microsoft). Develop comprehensive documentation for all the safety work, including metrics, evaluations, and progress tracking across multiple teams within OpenAI. Help with publishing and open sourcing safety
Jobs in United States
Ai Research Fellowship in United States
5,418 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai research fellowship jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Core Models team helps shape how OpenAI’s frontier models are built, measured, and launched. We work across Research, Engineering, Model Design, Data Science, and Product to turn advances in model capabilities into reliable, useful experiences for people. Our scope includes model planning and launches as well as building data flywheels, evaluations and measurement systems to ensure our models have strong capabilities and behavior. About the Role As a Product Manager for the Core Models team, you'll be at the forefront of defining and guiding the future of how our AI models work in real-world applications. You will connect user needs to model and systems decisions: how prompts are understood; how information is aggregated and made useful for training and evaluation data; and how capabilities move from research prototypes into the mainline model and launch stack. You will operate comfortably across research, infrastructure, and consumer product surfaces, creating clarity where ownership and technical boundaries are still emerging. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate user and product goals into clear model requirements, system architecture choices, and research priorities across query understanding, indexing, retrieval, ranking, tool boundaries, data, training, inference, and evaluation. Build closed learning loops that turn product usage, explicit feedback, and other user signals into datasets, evaluations, experiments, training priorities, and launch decisions. Define success across offline evaluations and online product metrics, balancing model quality, usefulness, latency, safety, reliability, and cost. Partner closely with post-training research, applied product engineering, Model Design, and Data Science to integrate capabilities into the mainline model stack. Create reusable platforms and operatin
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a
About the Team The Corporate Development team sits at the intersection of OpenAI’s strategy and the global AI ecosystem. Partnering closely with research, engineering, product, GTM and business leaders, we analyze markets, technologies, and companies to identify opportunities that accelerate OpenAI’s mission and long-term strategy. Bringing together expertise in strategy, M&A, and integration, the team leads the end-to-end process of sourcing, evaluating, negotiating, and integrating acquisitions and other strategic transactions. About the Role This is a high-visibility, strategically critical, cross-functional role at the intersection of M&A execution, enterprise go-to-market, and operational excellence. You will lead the critical phase from diligence through integration planning and post-close follow-through, define repeatable integration playbooks, and partner directly with senior stakeholders to ensure each transaction delivers intended value. The ideal candidate brings structured, hands-on integration experience, with particular expertise in enterprise / B2B transactions and a passion for AI, including building an AI-native corp dev platform. This role is based in our San Francisco HQ. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead integration planning and strategic execution for acquisitions from diligence through post-close follow-through, with clear owners, milestones, risks, and decision points. Translate deal rationale into practical integration goals and plans across product, engineering, sales, partnerships, customer success, people, finance, legal, security, systems, and communications. Partner with deal leads during evaluation and diligence to assess integration considerations, including tech stack, talent assessment, GTM alignment, transaction structure, operational readiness, and execution risk. Advise sponsors on how best to bring onboard acquisitions
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities Co-design future hardware for programmability and performance with our hardware vendors Assist hardware vendors in developing optimal kernels and add support for it in our compiler Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory h
About the Team OpenAI's Strategic Sourcing team helps the company scale responsibly, efficiently, and at speed. We partner with leaders across Engineering, Research, IT, Systems, Finance, Legal, and other teams to shape commercial strategies, negotiate critical agreements, and build resilient supplier ecosystems. Our work connects technical strategy, commercial judgment, financial discipline, and execution in support of OpenAI's mission. About the Role OpenAI is seeking a Strategic Technology Negotiations Lead to personally lead some of the company's most complex and consequential technology negotiations. This is not a people manager role. It is a senior strategic IC operator role for an expert negotiator who wants to remain close to the work and personally drive high-impact outcomes. We are especially interested in leaders who have managed teams and are intentionally seeking an individual contributor role where their impact comes through judgment, influence, and direct ownership. Rather than owning a fixed category, you will be deployed against high-priority opportunities where deal complexity, commercial stakes, executive visibility, or time pressure require exceptional negotiation leadership. Your initial focus will include data platforms and infrastructure, including data lake and lakehouse technologies, observability, and enterprise SaaS, with flexibility to work across other strategic technology areas. You will lead negotiations from strategy through execution, aligning decision-makers and driving agreements to closure. Many of these negotiations exist within broader supplier and partner ecosystems. You will look beyond the immediate transaction to account for interconnected cost, equity, revenue, partnership, risk, and long-term strategic implications. Success requires strong economics, sound judgment under pressure, executive credibility, and the ability to bring stakeholders with you through difficult decisions. This role is based in San Francisco, CA. We u
About the Team The Platform Systems team at OpenAI operates at the intersection of cutting-edge AI and large-scale distributed systems. We build the engineering and research infrastructure required to train OpenAI’s flagship models on some of the world’s largest, custom-built supercomputers. Our team develops core model training software and works deep in the stack - spanning collective communication, compute efficiency, parallelism strategies, fault tolerance, failure detection, and observability. The systems we build are foundational to OpenAI’s research velocity, enabling reliable, efficient training at frontier scale. We collaborate closely with researchers across the organization, continuously incorporating learnings from across OpenAI into the evolution of our training platform. About the Role As a Software Engineer, Platform Systems, you will design and build distributed systems that provide visibility into large-scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical to operating OpenAI’s training stack and is actively evolving to support new use cases and increasingly complex workloads. This role sits at the core of our training infrastructure, blending systems engineering, performance analysis, and large-scale debugging. In This Role, You Will Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high-throughput distributed systems Collaborate with systems, infrastructure, and research teams to evolve platform
About the team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing together world-class scientists, engineers, and business leaders to translate frontier research into real-world impact. Within OpenAI, the Go-to-Market organization helps customers understand, adopt, and scale our products across their businesses. The team includes Sales, Solutions, Support, Marketing, Partnerships, and Strategic Pursuits professionals who work together to bring the benefits of AI to organizations globally. About the role We are hiring a Senior Specialist Seller to join the Strategic Pursuits team and lead priority opportunities generated through the joint AWS and OpenAI co-sell motion. You will partner closely with Account Directors, AWS field teams, and cross-functional stakeholders to identify, shape, advance, and close high-value strategic enterprise engagements. This role is ideal for a senior seller with deep experience in enterprise technology, AI, cloud, or complex transformation sales motions. You should be comfortable operating in executive-facing environments, aligning multiple stakeholders, and turning customer priorities into commercially compelling and technically credible opportunities. This role is based in San Francisco, Seattle, or New York. We operate on a hybrid model of 3 days per week in office and offer relocation support. We are also open to remote candidates based in the United States. In this role, you will: Lead commercial execution of priority AWS/OpenAI co-sell opportunities across strategic enterprise accounts. Identify and qualify high-potential opportunities focused on GenAI adoption, frontier platform use cases, and enterprise-wide transformation. Partner with AWS and OpenAI field teams to align account strategy, messaging, stakeholder mapping, and execution plans. Develop customer-facing value propositions that connect OpenAI capabilities with business priorities and
About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Engineering Program Manager, you will help turn complex infrastructure strategy into executable programs across electrical, mechanical, controls, network, hardware, construction, commissioning, deployment, and operations workstreams. You will partner with research, hardware engineering, data center engineering, site development, supply chain, security, EHS, finance, legal, operations, and external delivery partners to bring OpenAI's infrastructure vision to life. About the Role We are looking for an Engineering Program Manager (EPM) to lead assigned infrastructure programs focused on production and non-production network integration, controls coordination, and the design and deployment of data hall or whitespace facilities. The EPM will support functional Directly Responsible Individuals (DRIs) across network, controls, structural, electrical, and mechanical disciplines. Key responsibilities include coordinating assigned workstreams and program controls, maintaining risks and interfaces, and supporting readiness within the network and data hall deployment track. The ideal candidate thrives on bringing structure to complex environments characterized by ambiguous technical requirements, large partner ecosystems, tight deadlines, and high operational stakes. This individual must be adept at keeping teams aligned on decisions, risks, dependencies, schedules, and readiness criteria, and escalating gaps or decision points when needed. Candidates should have a proven track record of managing technically challenging engineering programs across major lifecycle phases, including design, validation, procurement, construction, c
Other cities to consider
More places hiring for this role
Get new ai research fellowship jobs in United States by email
Daily job updates · Unsubscribe anytime