About the Team The Frontier Assurance team brings independent scrutiny into OpenAI’s safety decisions and helps the public understand and assess our safety work. We lead third-party assessments and safeguard testing for OpenAI’s flagship launches, pilot new assurance mechanisms such as embedded auditing, run our misalignment disclosure process, and incorporate independent expert input as evidence for critical safety decisions. About the Role As a Research Program Manager on the Frontier Assurance team, you will build programs that bring independent expertise into frontier AI safety decisions and make the evidence behind those decisions understandable to the public. You will lead external research partnerships and third-party assessments, coordinate public safety documentation, and develop new approaches to independent scrutiny and transparency. Working across research, engineering, product, policy, and communications, you will help ensure external findings inform concrete decisions and that our public explanations accurately reflect the evidence, limitations, and remaining uncertainty. We’re looking for people with deep experience in research partnerships and program management with technical and research teams. This role combines partnership management, cross-functional coordination, an understanding of AI safety research, alignment, and evaluations, and strong communication skills. You will work with researchers and engineers within OpenAI and across the external community to initiate projects, set ambitious goals and milestones, and drive execution across multiple teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and run third-party assessment programs for frontier models and safeguards, including independent evaluations, adversarial testing, and new approaches such as embedded auditing. Work with researchers and external part
Jobs in United States
Auditing in San Francisco
10 active opportunities · Updated September 2026
Showing
10 jobs
Explore current auditing jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make rigorous, evidence-based talent decisions and ensures that the systems supporting those decisions are valid, reliable, fair, and accountable. About the Role As a People Data Scientist focused on AI fairness and bias testing, you will help establish how OpenAI evaluates AI-assisted People systems and high-impact talent processes. You will design and conduct rigorous assessments to identify, measure, and mitigate potential bias across the lifecycle of models, agents, decision-support tools, and automated workflows. Your work will span the entire employee life-cycle, such as hiring, performance, promotion, employee development, workforce planning, etc. You will evaluate both technical systems and the broader human-AI decision processes in which they operate, examining not only model performance but also data quality, measurement validity, differential outcomes, human oversight, and unintended consequences. We’re looking for an experienced data scientist or applied researcher who can translate complex fairness questions into defensible evaluation strategies, scalable testing infrastructure, and clear recommendations for technical teams and senior leaders. This role is preferred to be based in San Francisco, CA. In this role, you will: Define and lead fairness and bias-testing strategies for AI-assisted People processes, models, agents, and decision-support systems from development through deployment and ongoing monitoring. Design rigorous algorithmic audits and validation studies, including adverse-impact analysis, subgroup and intersectional evaluation, error-rate analysis, calibration, measurement invariance, reliability, criterion-related validity, and sensitivity testing. Identify the appropriate fairness criteria for each use case, evaluate tradeoffs among competing definitions
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Integrations Operations Engineering (IOE) team strengthens Plaid's network with financial institutions by directly resolving integration issues to increase the reliability and quality of data, and by building new integrations to grow our financial network. We sit at the nexus of engineering, a deep understanding of Plaid's products and customers, and direct financial-institution relationships: we write and ship the code that keeps connectivity healthy, we use data to focus on the issues with the greatest customer impact, and we work directly with data partners (the financial institutions and platforms themselves) to resolve the problems that can't be fixed from our side alone. The team plays a mission-critical role in ensuring industry-leading connectivity so our customers can meet ever-expanding financial-services use cases and reach as many users as possible. What You'll Do Investigate and resolve the highest-impact integration issues by writing maintainable, tested code and deploying it to production, then monitor for regression or degradation after your changes ship. Prioritize by customer impact. The team runs a business-value-based prioritization model that automatically
About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About the Team The People Operations team owns the systems, processes, and programs that make the employee lifecycle run — from day one through every transition in between. We believe that clean data, simple workflows, and well-designed technology are what let the rest of the People & Talent function do its best work. About the Role We're looking for a People Operations Manager who can look at a messy, ambiguous problem and turn it into a clear plan — then execute it. Day-to-day, you'll own our HR technology stack, lead process transformation across the employee lifecycle, and serve as the go-to partner for Payroll, IT, Legal, and the broader People & Talent team when something needs to be built, fixed, or rethought. You'll work directly with vendors, HRBPs, and leadership to make sure our operations are scalable, compliant, and actually simple to use. In the first year, success means moving us from where we are to where we should be: better-configured systems, documented processes, and a team that trusts what's in our tools. Responsibilities Lead HR operations programs end-to-end — define scope, owners, milestones, risks, and success measures, often starting from an undefined brief. Serve as the primary functional owner for core People technology platforms (e.g., HRIS, ATS, LMS, performance tools), including configuration, upgrades, and feature evaluation. Lead technology deployments, platform conversions, and module rollouts, coordinating across IT, HRIS, and vendor teams. Research and benchmark HR technology trends to build and maintain a People tech roadmap that supports scale. Apply AI and e
About the Team The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability , which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor). Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research.We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable. About the Role We’re looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work. As a researcher on the Alignment team, you will design and run experiments that improve our understanding of model monitorability. You will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make those questions measurable, and help translate findings into practical oversight and training recommendations. You may also help develop new monitoring models or methods and apply them to OpenAI’s largest training runs. This role is especially well
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
About the team Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. Mitigating the frontier risks resulting from these capabilities is paramount to OpenAI’s ability to continue deploying models safely. The Preparedness team is dedicated to addressing these critical risks. Our work includes: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misalignment safeguards, alignment tools, and on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the role Preparedness is hiring strong technical executors to support preparations for accelerated AI development, which may culminate in recursive self-improvement. This work relies on anticipating misalignment risks that might exist in the future, but might not exist now; so it’s especially important that people in this role are tasteful and strategic. The role is wide-ranging, covering any mitigation for loss of control risk, spanning the design and implementation of better pre-deployment risk-assessment , control measures , RSI-relevant training interventions, and turning one’s technical work into established institutional practices and external-facing communications. Below is a subset of our focus areas: Scalable oversight: Establishing practices for model misbehavior monitoring and oversight which remain effective in superhuman model capability regimes, with a focus on bridging from today’s monitoring approaches to future-proof ones. Automated auditing: As model capabilities increase, we’ll increasingly rely on automated approaches for finding the most severe forms of model misalignments. We’ll both need to s
About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev
About the Team The Future of Computing Research team is an Applied Research team within the Consumer Devices group focused on developing new methods and models as we advance forward in our mission of building AGI that benefits all of humanity. As a Software Engineer on the Future of Computing Research team, you will work together with both the best ML researchers in the world and the greatest design talent of our generation to push the frontier of model capabilities. About the Role We are looking for a Software Engineer to join our team to build tools and services that enable AI research, evaluation, and data generation workflows. The best work in this role will start with an ambiguous design question and turn it into working research systems. You will work closely with researchers, designers, and engineers to build the evaluation systems, synthetic data generation pipelines, review tools, and supporting platform services. The goal is to make these workflows easier to create, run, and trust without requiring bespoke engineering support for each new design concept. You will help ensure that research artifacts have a clear lifecycle, runs are reproducible and observable, and results provide useful evidence for product and model-training decisions while the underlying systems remain reliable and reusable. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Build web applications, APIs, data models, and backend services for AI research workflows. Build tools to author and manage evaluation tasks, rubrics, graders, suites, and rollout configurations, including workflows for publishing, versioning, auditing, and sharing research artifacts. Automate evaluation runs and generate useful reports for design, research, and engineering teams. Support synthetic data generation workflows for multimodal and conversational research, including tools that comb
Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity We are seeking a highly analytical and detail-oriented Sales Compensation Analyst to own the administration and reconciliation of our sales incentive programs. You will be the primary engine behind our commission cycles, ensuring all employees on variable compensation plans are paid accurately and on time. Your mission is to bring absolute data integrity to our compensation process. You will bridge the gap between CRM data and payroll outputs, resolve discrepancies, and serve as a trusted partner to both Finance and the Sales organization. What You’ll Do Monthly audit of CRM bookings vs. commission workbooks and tools to verify deal values, credits, and contract terms. Administrative Workflow: Manage the end to end commission lifecycle, from data extraction to final payroll delivery. Dispute Resolution: Serve as the first point of contact for commission inquiries, investigating and resolving rep payout questions. Data Hygiene & Auditing: Conduct regular audits of CRM opportunity stages and account ownership to ensure the commission system aligns with current business reality. System Maintenance: Maintain a
Other cities to consider
More places hiring for this role
Get new auditing jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime