About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system
Jobs in United States
Ai Tech Lead in United States
5,082 active opportunities · Updated October 2026
Showing
15 jobs
Explore current ai tech lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role Our technologies support some of the most important and impactful work in the world, including our strategic and high-impact customers in the public sector. As a Forward Deployed Security Engineer (FDSecE) you will be responsible for securing these novel applications of OpenAI’s technology. We’re looking for motivated, tenacious, and curious people who will work closely with engineering teams to ensure our infrastructure deployments are highly secure against our adversaries. As an FDSecE, you will embed directly throughout the lifecycle, working on-site and being hands-on to ensure the overall security of these deployments from design to production and through ongoing operations. This role is preferred to be based in Washington DC but may consider remote work. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Travel to and working from customer sites is required for this role. In this role, you will: Deeply embed with our most strategic public sector customers to implement and maintain robust security controls. Be a design and technical thought partner by leveraging security expertise on protective controls including access controls, authentication, encryption, network, and system security. Collaborate closely with teammates, cross-functional teams, customers, and service providers to achieve security and compliance goals. Ensure continuity of critical security and monitoring c
About the Team OpenAI’s mission is to ensure the responsible and widespread adoption of artificial intelligence. In support of that mission, the Marketing team helps deeply understand customer audiences and market dynamics, influence the development of the right products, build sustainable and customer-aligned monetization models, and drive awareness, adoption, and usage across OpenAI’s products and platform. We take a data-driven approach to understand markets, develop monetization strategies, and uncover customer needs that shape product strategy and messaging. We partner closely with Sales, Partnerships, Product, Engineering, Research, Comms, and Design to deliver a cohesive end-to-end customer experience and lead go-to-market efforts for new product launches across channels. About the Role We are hiring a Head of Enterprise Ads Marketing to define and scale OpenAI’s global marketing strategy for enterprise advertisers, agencies, and strategic brand partners. This leader will own how advertisers understand, evaluate, adopt, and grow with OpenAI’s advertising solutions. This is a highly strategic and hands-on leadership role where you will shape the narratives that position OpenAI as a premium advertising platform for the world’s marketers while building the content, thought leadership, events, and demand programs that drive pipeline and long-term strategic growth. You will partner closely with Sales, Partnerships, Product, Finance, Legal, Communications, Research, and Data Science to help establish OpenAI as a trusted leader in the future of AI-powered advertising. This role is ideally based in San Francisco, CA, with a hybrid schedule of three days per week in the office. In This Role, You Will: Define the end-to-end enterprise advertiser marketing strategy across awareness, consideration, pipeline creation, adoption, expansion, and long-term growth. Develop OpenAI’s enterprise advertising narratives, positioning, and messaging architecture across AI-powered mar
About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Strategic Sourcing & Procurement function plays a critical role in enabling OpenAI to deliver impact across research, product development, technology infrastructure, and services by helping the company scale responsibly, securely, and with strong commercial discipline. Our work sits at the intersection of innovation and execution. We partner closely with teams across OpenAI to translate rapidly evolving business needs into scalable, compliant, and economically sound external partnerships. As OpenAI continues to grow at pace, services sourcing is becoming increasingly strategic across the company. Every business unit relies on external service providers in different ways — to extend capacity, access specialized expertise, support operations, and accelerate execution. Done well, Procurement becomes a source of trust and momentum, helping OpenAI move faster with the right partners, stronger commercial outcomes, and the right level of protection. About the Role We are seeking an experienced Strategic Sourcing (GTM) Leader to lead strategic sourcing and commercial enablement for OpenAI’s Go-to-Market organization across B2B and B2C channels. You will manage substantial and rapidly growing spend while shaping sourcing strategies and scalable commercial pathways across Media, Creative, Production, Influencer, Agency, Sponsorships, Analytics, Communications, and Event suppliers in support of high-impact global initiatives. You’ll help evolve our GTM procurement function from reactive deal support into a speed-enabling, scalable commercial engine that delivers cost efficiency, launch readiness, and strong governance in a fast-moving environment. In this role, you will: Develop and execute sourcing strategies across GTM, Brand, Global Affairs, Events, Growth, and Partnership activities—spanning both B2B and B2C channels—that align with our mission and b
About the Team At OpenAI, we’re building the connective tissue between our mission and our people. People Innovation Labs is a fast-moving engineering team embedded in the People organization, focused on rethinking how we find and retain the best talent and empower everyone to do their best work. From recruiting to culture, we’re designing systems that give our People Team a significant edge by infusing OpenAI’s models and first-principles thinking into every aspect of our work. Our projects range from greenfield 0-1 products like OpenHouse (our internal knowledge hub) to AI-powered automations and scalable recruiting tools. We’re defining the future of work at OpenAI, creating a blueprint for how AI can supercharge productivity, culture, and innovation. About the Role We’re seeking a Data Engineer to build data-intensive systems that will power People Innovation Labs’ internal products and enable the People Analytics function to do their best work. These data pipelines are crucial for our build-out of people products backed by business systems of record and for ongoing people data analytics. One example of an employee-facing product you’ll help us build is OpenHouse, which serves as a culture and communication hub and an organization-wide front door into all other aspects of People Innovation Labs’ work. OpenHouse and other products in our portfolio are built by full stack product engineers who are deeply curious about culture, recruiting and people development, and want to know everything from the business strategy and metrics down through the code that gets us there. In this role, you will work with People Innovation Labs leadership and software engineers and the People Analytics team to build the data systems that enable this work. In this role, you will: Design, build and manage people data pipelines, ensuring all data is seamlessly integrated into our Databricks warehouse. Develop canonical datasets to track key people metrics and People Innovation Labs produc
About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. The team partners closely with research and product teams across the company, and conducts research as a final step to prepare for real world deployment to millions of users, ensuring that our models are safe, efficient, and reliable. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve model capability and performance. Collaborate closely with the other research and product teams, allowing customers to optimize their own models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of relevant models, and building evaluations for model capability improvement. Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and
About the Team OpenAI's Research Data Team exists to accelerate the evaluation, safety and capabilities of our models and products. Made up of technical operators and software engineers, we design the methods in which we acquire and create data. About the Role As a Research Program Manager (RPM), Data Acquisition, you will partner with research, engineering, and operations to design and implement pragmatic solutions for acquiring data. You will be a key interface between our research roadmap and external data offerings. This role is based in our San Francisco HQ and will be part of a team of RPMs pushing the frontier of data acquisition. In this role, you will: Partner deeply with research: Work with researchers to scope data needs, define success criteria, and translate priorities into clear execution plans. Shape the data acquisition pipeline: Identify, evaluate, and advance high impact data opportunities - balancing research value, feasibility, quality, and responsible execution. Unblock yourself: Move work forward even when the path is unclear — using technical judgement, creative problem solving, and scrappy execution to make progress while longer-term solutions are still forming. Build lightweight systems and visibility: Use SQL, Python, dashboards, and simple tooling to track performance, quality, and blockers. Drive technical roadmaps: Collaborate with engineers to enhance data platforms, resolve blockers, and ensure security best practices such as access management. Scale your impact: Equip vendors and internal teams with the context, standards, and operating rhythms needed to focus on the most important problems. You’ll thrive in this role if you: Are proficient in SQL and Python for analysing datasets, querying databases, building dashboards, and generating actionable insights. Are comfortable using APIs, automation, and AI tools such as Codex to accelerate workflows, remove manual overhead, and upskill quickly in unfamiliar technical areas.Experience sou
About the Team Frontier Systems Foundations, part of Compute Foundations at OpenAI, builds the systems software foundation that turns new compute infrastructure into reliable, usable capacity for frontier model training. Our mission is to make some of the world's largest GPU clusters work reliably for frontier training. We bring new platforms and clusters online, safely maintain installed fleets, and partner with hardware, infrastructure, and research teams to resolve the system-level issues that keep jobs from running. That means building and maintaining the software closest to the machine: Linux and Ubuntu operating-system images, kernels and modules, drivers, packages and repositories, disks and boot configuration, firmware integration, provisioning, and system-level validation. We make these components reproducible, compatible, and safe to operate across heterogeneous fleets. About the Role We are looking for systems software engineers with deep Linux and host-systems experience to build, qualify, and maintain the operating-system foundation for OpenAI's frontier compute fleet. Relevant backgrounds include kernel and module development, Linux distribution or image engineering, package management, firmware and driver integration, disks and boot, and bare-metal provisioning. You'll work closely with hardware engineers, vendors, and infrastructure teams to bring up new platforms, integrate system components, and debug failures across firmware, disks, boot, operating systems, kernels, drivers, and workload interactions. Your work will directly influence how quickly new capacity becomes usable and how reliably large GPU fleets operate. You should be comfortable writing and maintaining production-quality systems software and automation, but we do not expect expertise across every layer. This is an opportunity to go deep on challenging systems problems while building the image, package, qualification, and recovery paths that power the next generation of frontier models
About the Team The IT and Security organization builds the systems, data foundations, and automation that help OpenAI operate securely and reliably at scale. We support critical domains across identity, access, infrastructure security, enterprise systems, and internal productivity. As OpenAI grows, audit readiness and control assurance increasingly depend on reliable data: accurate system inventories, access populations, change records, configuration state, exception signals, and evidence generated directly from source systems. Our goal is to move beyond manual evidence collection and build scalable data products, automated validation, and continuous control monitoring that make security and IT controls measurable, repeatable, and defensible. About the Role We are looking for an IT Controls Data Engineer to build the data infrastructure that powers audit readiness, IT controls, evidence automation, and continuous control monitoring. In this role, you will design and maintain the pipelines, datasets, models, validation logic, dashboards, and evidence exports that make IT controls measurable, repeatable, and defensible. You will work across Security, IT, Infrastructure, Engineering, Finance Risk Management, and auditors to turn complex system behavior into reliable control data products. This is a technical builder role. The ideal candidate is strong in data engineering and analytics engineering, comfortable working with enterprise and security system data, and able to explain data lineage, source-system behavior, and control logic clearly to technical and audit stakeholders. You’ll be responsible for Building reliable data pipelines, models, and datasets for IT controls, including access, identity, configuration, change, ticketing, exception, and evidence data. Creating data quality, lineage, reconciliation, and completeness checks that make control data defensible for SOX and other audit use cases. Designing automated evidence generation workflows that produce compl
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Artifacts, you will train frontier models to create polished, useful work products: documents, spreadsheets, slide decks, dashboards, reports, analyses, and other interactive or editable artifacts. You will help teach our models to move from a vague user goal to a finished artifact with strong structure, visual taste, domain judgment, correctness, and low latency. This work will require owning improvements across our post-training stack, including RL, data pipelines, graders, reward signals, evals, and behavioral analysis. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you will: Design and run experiments that improve agentic model behavior for complex so
About the Team The Storage Infrastructure team builds and operates the storage foundation behind OpenAI’s most demanding workloads. We work directly with research to design storage systems for rapidly evolving experiments, while also powering production at scale. We own the platform end to end: backend systems, user-facing services and APIs, and the control planes that manage how data is placed, moved, and retained over time. Our stack spans cloud and in-house object stores across very different workload profiles, from GPU-attached systems to dedicated storage hardware. We also build the federation layer that unifies these backends behind a simple interface and routes each workload to the right storage solution. About the Role You will help build the storage platform that powers OpenAI’s research and production systems. This is a hands-on infrastructure role for engineers who want to work on deeply technical systems at scale and own them in production. You’ll work across object storage, cross-region data movement, lifecycle management, and the federation layer that provides a unified interface across multiple backends. Much of our stack runs on Kubernetes, and we primarily build services in Rust. In this role, you will: Build and operate storage services that underpin OpenAI’s research infrastructure Develop object storage systems across cloud and in-house environments Build systems for cross-region data movement, replication, and recovery Design lifecycle management capabilities that keep data durable, available, and cost-effective Evolve the federation layer that unifies multiple backend systems behind a simple interface Improve performance, reliability, and operational excellence across the platform Collaborate closely with researchers and infrastructure teams to support rapidly evolving workloads You might thrive in this role if you: Have experience building or operating distributed systems in production Have worked on storage infrastructure, object stores, dist
About the Team The Business Data Science team uses data and analytics to optimize business performance, drive growth, and foster meaningful partnerships, with the goal of ensuring the sustained and impactful expansion of OpenAI's initiatives to maximize the benefits of AGI for all of humanity. We partner with Sales (GTM), Marketing, Partnerships, Support, Finance, Product, and Growth. About the Role As a member of our Business Data Science team, you will help build a data-driven culture around insight generation, decision making, and strategy at OpenAI. This role is focused on driving customer success within our business products (ChatGPT Team, ChatGPT Enterprise, and API). You will work on projects such as identifying opportunities for interventions within a customer lifecycle to drive activation & onboarding, identifying target audiences for new feature launches, and measuring the efficacy of emails, events, and other interventions to drive ongoing engagement with our products. This role is based in San Francisco, CA or New York, NY. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Embed with our Customer Success organization as a trusted partner, uncovering new ways to drive customer adoption and engagement of our business products. Establish key metrics, run experiments, and perform analysis to help us understand the incrementality of our efforts to drive adoption/engagement. Proactively surface insights and opportunities to drive engagement and growth. Build tools and systems for stakeholders to self-serve routine data and insights freeing up time to work on more leveraged analyses. Become an expert in OpenAI’s data and systems. Through partnership with Data Eng, Finance and other business teams, you will self-serve all the underlying data for our business and derive insights from them. Partner with other data scientists across the company to share knowledge and continually
About the Team Codex is OpenAI’s first-party developer product focused on agentic software engineering. We’re building tools that help engineers design, write, test, and ship code faster—safely and at scale. We partner tightly with research and product to translate model advances into tangible developer productivity. About the Role As a Data Scientist on Codex, you will measure and accelerate product-market fit for AI developer tools. You’ll define what “developer productivity” means for our product, run experiments on new coding models and UX, and pinpoint where the model helps or hurts across languages and tasks. Your insights will directly shape how an entire industry builds software. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Codex product team to discover opportunities that improve developer outcomes and growth Design and interpret A/B tests and staged rollouts of new coding models and product features Define and operationalize metrics such as suggestion acceptance, edit distance, compile/test pass rates, task completion, latency, and session productivity Build dashboards and analyses that help the team self-serve answers to product questions (by language, framework, repo size, task type) Diagnose failure modes and partner with Research on targeted improvements (model quality signals, user feedback, evals) You might thrive in this role if you have 5+ years in a quantitative role at a developer-facing or high-growth product Fluency in SQL and Python; comfort with experiment design and causal inference Experience defining product metrics tied to user value Ability to communicate clearly with PM, Eng, and Design—and to influence product direction You could be an especially great fit if you have Strong programming background; ability to prototype, run simulations, and reason about code quality Familiarity with IDE/extensi
About the Team The Interpretability team studies internal representations of deep learning models. We are interested in using representations to understand model behavior, and in engineering models to have more understandable representations. We are particularly interested in applying our understanding to ensure the safety of powerful AI systems. Our working style is collaborative and curiosity-driven. About the Role OpenAI is seeking a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, in close collaboration with a highly motivated team. You will play a critical role in helping OpenAI ensure future models remain safe even as they grow in capability. This will make a significant impact on our goal of building and deploying safe AGI. In this role, you will: Develop and publish research on techniques for understanding representations of deep networks. Engineer infrastructure for studying model internals at scale. Collaborate across teams to work on projects that OpenAI is uniquely suited to pursue. Guide research directions toward demonstrable usefulness and/or long-term scalability. You might thrive in this role if you: Are excited about OpenAI’s mission of ensuring AGI benefits all of humanity, and are aligned with OpenAI’s charter . Show enthusiasm for long-term AI safety, and have thought deeply about technical paths to safe AGI. Bring experience in the field of AI safety, mechanistic interpretability, or spiritually related disciplines. Hold a Ph.D. or have research experience in computer science, machine learning, or a related field. Thrive in environments involving large-scale AI systems, and are excited to make use of OpenAI’s unique resources in this area. Possess 2+ years of research engineering experience and proficiency in Python or similar languages. Are deeply curious. About OpenA
About the Team Our infrastructure team helps deliver OpenAI’s most capable models and products to the world by scaling infrastructure and turning demand into useful FLOPS. We collaborate across research, engineering, design, and business to turn cutting-edge AI advancements into impactful, real-world applications. Our team ensures the right compute is available—at the right time and place—to support some of the world’s most demanding workloads. We empower all of OpenAI’s products and research by scaling the infrastructure behind them. Our work makes it possible to launch new models and products reliably and at scale. About the Role As a Data Scientist on the Infra team, you will play a key role in shaping how we scale the infrastructure that powers OpenAI’s products and research. This is critical as we operate one of the largest and most advanced compute fleets in the world, supporting millions of users and businesses globally. We focus on aligning infrastructure measurement, planning, scaling, allocation, and efficiency to drive measurable impact across the company. You should expect to guide the definition of foundational datasets for infrastructure resources, develop metrics that inform key decisions, build forecasting and optimization models, and establish source of truth dashboards and analyses that enable teams to understand and improve infra usage. Most importantly, you should expect to be a core partner to engineering, research, and product teams in shaping the infrastructure that powers everything OpenAI builds. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain foundational datasets and metrics that reflect infrastructure usage, efficiency, and scaling. Develop forecasting and optimization models to support infra planning and resource allocation. Partner with engineering, research, and product teams to shape infrast
Other cities to consider
More places hiring for this role
Get new ai tech lead jobs in United States by email
Daily job updates · Unsubscribe anytime