About the Team At OpenAI, the User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving company priorities. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are looking for a senior program manager to build the safety, quality, and risk operations supporting a new category of consumer devices. You will translate ambiguous product risks and evolving requirements into practical operating models, workflows, escalation paths, launch-readiness plans, and cross-functional decision-making. This is a foundational role: the systems you build will shape how OpenAI launches, monitors, and improves a new category of consumer devices safely at scale. You will help establish how potential safety incidents, product-quality concerns, sensitive customer escalations, privacy-sensitive issues, and other emerging device risks are identified, investigated, resolved, and incorporated into product and operational improvements. You will turn incomplete requirements into practical workflows, decision rights, launch plans, quality controls, measurement, and durable ownership. The role centers on program building, operational judgment, and execution. We welcome candidates from product safety, quality assurance, regulatory operations, technical program management, healthcare, medical devices, aerospace, consumer technology, and other environments involving complex products or regulated risks. Direct consumer-hardware experience is helpful but not required. The strongest candidates learn unfam
Jobs in United States
Investigator in San Francisco
57 active opportunities · Updated September 2026
Showing
15 jobs
Explore current investigator jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The OpenAI Audit Team is on a mission to build the future of internal audit from the ground up. Our ambition will be powered by a high-energy, technically exceptional team with the judgment, intellectual curiosity, and creativity to harness the latest advances in AI and design a truly next-generation audit function. As AI reshapes how work is performed across the enterprise, Internal Audit will use AI, automation, and data analytics to identify and assess the most significant and emerging risks across the business, including technology, cybersecurity, finance, compliance, operations, and data. We will build trusted partnerships at every level—from the Board of Directors and senior leadership to the teams delivering on OpenAI’s mission every day. We will operate as both an independent assurance provider and a trusted advisor, bringing an objective and pragmatic perspective to critical decisions. By engaging closely with management while preserving our independence, we will help the business innovate responsibly, move with confidence, and manage risk without creating unnecessary barriers. About the Role As the Finance & Operations Audit Leader, you will help shape the strategy, methodology, technology, and culture of a new audit function. You will lead complex audits, forensic and investigative reviews, and advisory work across financial reporting, accounting, finance, and business operations, while addressing related compliance risks. You will advise leaders on practical ways to strengthen governance, execution, accountability, and risk management. You are an experienced, hands-on professional who combines deep finance and accounting knowledge with strong audit, risk, and business judgment. Your work will span across finance & accounting areas including treasury, tax, revenue, financial planning, business operations, and enterprise governance, with opportunities to apply forensic accounting techniques to complex transactions, anomalies, and pot
About the Team OpenAI Finance ensures the organization is positioned for long-term success as we pursue our mission. The Revenue team plays a critical role in enabling OpenAI to scale its commercial offerings—overseeing billing operations, deal desk, revenue systems, and revenue accounting. We work cross-functionally with Technical Revenue, Finance Data, and Revenue Systems teams to support complex commercial arrangements, improve operational efficiency, and maintain financial integrity. About the Role As a Revenue Accounting Manager, you will own revenue accounting processes for consumption and usage-based revenue recognition while helping implement and maintain the systems, data flows, and accounting rules that support accurate financial reporting. You will serve as a key execution partner on onboarding new revenue streams, Fusion Accounting Hub rule updates, translating accounting requirements into expected journal entries, source-to-general-ledger mappings, user acceptance testing, data validation, and controlled operating processes. We’re looking for a hands-on revenue accounting owner who combines strong close discipline with systems and data fluency, independently coordinates cross-functional implementation work, and strengthens the accounting infrastructure supporting OpenAI’s growth. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the monthly close for usage-based revenue and payment processor accounting, including journal entries, reconciliations, variance analysis, controls, and supporting documentation. Prepare and review revenue-related journal entries while understanding the underlying transaction lifecycle, accounting methodology, billing arrangements, source data, and expected financial reporting outcomes. Perform reconciliations for key revenue accounts, investigate discrepancies, and drive issues to resolution. Own flux
About the Team GTM Growth Engineering builds AI-native products that help OpenAI's go-to-market and B2B marketing organizations scale with greater speed, intelligence, and operational effectiveness. We apply OpenAI models to real business workflows and build the systems that make those applications useful and dependable: customer context, agent behavior, feedback, evaluation, experimentation, and appropriate human oversight. Our work brings together software engineering, applied AI, product, data, and GTM operations. We measure success through the quality of customer engagement, pipeline, conversion, and the effectiveness of our sales and marketing teams. About the Role We're looking for an Applied AI Engineer to build production systems that help AI-powered go-to-market workflows improve over time. You will connect agent behavior, customer and operator feedback, evaluation, experimentation, and business outcomes to make these systems more effective, reliable, and responsive to evolving customer needs. This is a deeply technical, cross-functional role with end-to-end ownership of the agent improvement loop: understand production behavior, identify failure modes, improve how the system decides or acts, and validate the resulting impact. You will partner with Engineering, Product, Data Science, Sales, and B2B Marketing to turn real-world signals into safer, more effective agent behavior and measurable improvements in customer engagement, conversion, qualified pipeline, and team productivity. In this role, you will: Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes. Instrument agent workflows so model interactions, tool use, decisions, failures, human edits, and downstream outcomes can be understood in context. Define meaningful quality standards, representative evaluation datasets, regression coverage, and production monitoring for real GTM workflows. Investigate why a
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Secure Manufacturing & Stealth (SMS) team protects OpenAI’s most sensitive product, manufacturing, launch, and partnership work from theft, leakage, espionage, and unauthorized disclosure. We operate across physical, digital, human, and third-party domains, building practical protections that allow teams to move quickly while preserving strict need-to-know controls. About the Role We are seeking a senior Secure Manufacturing & Stealth Partner to serve as the dedicated SMS owner for Marketing. This person will build trusted relationships across Marketing, understand launches and sensitive information flows, identify secrecy risks early, and embed practical controls into planning, creative production, events, agencies, vendors, media, and executive communications. The role owns the Marketing relationship while coordinating shared SMS capabilities when investigative, prototype-handling, regional, or other specialist support is needed. In this role you will: Serve as the primary SMS Partner and trusted adviser to Marketing, building a deep understanding of its priorities, planning cycles, launches, events, external partners, and sensitive information flows. Identify secrecy and information-exposure risks early, then translate SMS requirements into practical controls that protect sensitive work without unnecessarily slowing execution. Build, document, and socialize durable operating mechanisms for compartmentalization, need-to-know access, agencies and vendors, sensitive creative production, events, demonstrations, launch preparation, codenames, watermarking, and password controls. Assess and approve Marketing systems and information workflows, identify gaps or duplication, establish secure handling requirements, and advise AI-native Marketing initiatives on permissions, data governance, and operational tracking. Coordinate
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching
About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving areas of work. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are seeking a Device Safety & Risk Operations Specialist to build the safety operating model for a new category of consumer hardware. This is a senior individual-contributor role for someone who can turn emerging product risks and incomplete requirements into practical workflows, controls, launch plans, and durable systems. You will define how product-safety incidents, critical escalations, regulated cases, and privacy-sensitive issues should be identified, investigated, escalated, resolved, and learned from. You will also establish operational requirements for case management, data access, decision logging, quality assurance, monitoring, and cross-functional response. You will stand up priority workflows through launch and early operations, then help transition them into durable homes across USRO and partner teams. The right person combines deep operational judgment with strong technical and hardware product fluency. They can move from executive-level risk framing to detailed workflow design, tabletop exercises, launch readiness, frontline guidance, and post-launch improvement. Location / work model: San Francisco, CA; hybrid, 3 days/week in-office. Please note: This role may involve exposure to sensitive or concerning material. Strong discretion, judgment, and resilience are essential. In This Role, You Will:
About the Team The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability , which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor). Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research.We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable. About the Role We’re looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work. As a researcher on the Alignment team, you will design and run experiments that improve our understanding of model monitorability. You will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make those questions measurable, and help translate findings into practical oversight and training recommendations. You may also help develop new monitoring models or methods and apply them to OpenAI’s largest training runs. This role is especially well
About the Team At OpenAI, our User Safety & Risk Operations (USRO) team helps protect our products and users from abuse, fraud, safety risks, and other forms of misuse. We operate at the front line of real-world safety and risk management, translating user and operational signals into timely decisions, effective interventions, and improvements to our systems. This role sits on a team focused on building operational capacity for new, ambiguous, and fast-moving areas of work. The team defines what needs to be built, creates the operating model to support it, and works with partner teams to make the work scalable and durable over time. About the Role We are seeking a Device Safety & Risk Operations Specialist to build the safety operating model for a new category of consumer hardware. This is a senior individual-contributor role for someone who can turn emerging product risks and incomplete requirements into practical workflows, controls, launch plans, and durable systems. You will define how product-safety incidents, critical escalations, regulated cases, and privacy-sensitive issues should be identified, investigated, escalated, resolved, and learned from. You will also establish operational requirements for case management, data access, decision logging, quality assurance, monitoring, and cross-functional response. You will stand up priority workflows through launch and early operations, then help transition them into durable homes across USRO and partner teams. The right person combines deep operational judgment with strong technical and hardware product fluency. They can move from executive-level risk framing to detailed workflow design, tabletop exercises, launch readiness, frontline guidance, and post-launch improvement. Location / work model: San Francisco, CA; hybrid, 3 days/week in-office. Please note: This role may involve exposure to sensitive or concerning material. Strong discretion, judgment, and resilience are essential. In This Role, You Will:
About the team The Applied team safely brings OpenAI's technology to the world. We released ChatGPT; Plugins; DALL·E; and the APIs for GPT-5, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. Our customers build fast-growing businesses around our APIs, which power product features that were never before possible. ChatGPT is a prime example of what is currently possible. We simultaneously ensure that our powerful tools are used responsibly. Safe deployment is more important to us than unfettered growth. The Fraud Engineering team works within our Applied Engineering organization identifying and responding to fraudsters on our platform. We are looking for a software engineer with anti fraud & abuse experience to help architect and build our next-generation anti-fraud systems. About the role The Scaled Abuse team protects OpenAI’s products and customers by detecting, preventing, and responding to fraudulent and abusive behavior at scale. We build and operate the backend and data systems that power real-time detection, investigation workflows, and enforcement — balancing strong protections with a great user experience as the platform grows. Our work sits at the intersection of engineering and abuse expertise: we partner closely with Trust & Safety, Security, and Product to understand emerging attack patterns, translate messy signals into clear system behavior, and continuously harden our defenses. The problems are dynamic and ambiguous by default, so we value engineers who can quickly dive into an unfamiliar codebase, develop strong intuition about how it works end-to-end, and propose pragmatic improvements that make the entire stack more resilient. In this role, you will: Design and build systems for fraud detection and remediation while balancing fraud loss, cost of implementation, and customer experience Work closely with finance, security, product, research, and trust & safety ope
About the Team: OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role Overview We are seeking a Package Reliability Engineer to lead reliability engineering for advanced packages used in high-performance AI and computing systems. The primary focus of this role is to assess package level mechanical and thermal reliability risks and apply thermal and mechanical modeling to optimize package design, material selection, and assembly processes. The engineer will also develop reliability test plans with external partners, identify failure mechanisms, perform root-cause analysis, and recommend practical corrective actions. In this role, you will assess package reliability risks from early architecture development through product qualification and high-volume manufacturing. You will work closely with package design, silicon design, system engineering, manufacturing, and ASIC partners to predict package behavior, develop qualification strategies, resolve reliability issues, and improve overall package robustness and lifetime. In this role you will: Lead reliability test plan and assessments for advanced HPC packages, including risk identification, potential failure-mechanism analysis, root-cause investigation, mitigation planning, and corrective-action development. Drive reliability-focused package design optimization based on thermo-mechanical modeling to improve package reliability, power integrity, thermal performance, mechanical robustness, and platform scalability. Develop, validate, and apply package reliability models and lifetime-prediction
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system
About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati
Other cities to consider
More places hiring for this role
Get new investigator jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime