About the Team The Business Data Science team uses data and analytics to optimize business performance, drive growth, and foster meaningful partnerships, with the goal of ensuring the sustained and impactful expansion of OpenAI's initiatives to maximize the benefits of AGI for all of humanity. We partner with Sales (GTM), Marketing, Partnerships, Support, Finance, Product, and Growth. About the Role As a member of our Business Data Science team, you will help build a data-driven culture around insight generation, decision making, and strategy at OpenAI. This role is focused on driving customer success within our business products (ChatGPT Team, ChatGPT Enterprise, and API). You will work on projects such as identifying opportunities for interventions within a customer lifecycle to drive activation & onboarding, identifying target audiences for new feature launches, and measuring the efficacy of emails, events, and other interventions to drive ongoing engagement with our products. This role is based in San Francisco, CA or New York, NY. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Embed with our Customer Success organization as a trusted partner, uncovering new ways to drive customer adoption and engagement of our business products. Establish key metrics, run experiments, and perform analysis to help us understand the incrementality of our efforts to drive adoption/engagement. Proactively surface insights and opportunities to drive engagement and growth. Build tools and systems for stakeholders to self-serve routine data and insights freeing up time to work on more leveraged analyses. Become an expert in OpenAI’s data and systems. Through partnership with Data Eng, Finance and other business teams, you will self-serve all the underlying data for our business and derive insights from them. Partner with other data scientists across the company to share knowledge and continually
Jobs in United States
Scientist in San Francisco
98 active opportunities · Updated October 2026
Showing
15 jobs
Explore current scientist jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team Codex is OpenAI’s first-party developer product focused on agentic software engineering. We’re building tools that help engineers design, write, test, and ship code faster—safely and at scale. We partner tightly with research and product to translate model advances into tangible developer productivity. About the Role As a Data Scientist on Codex, you will measure and accelerate product-market fit for AI developer tools. You’ll define what “developer productivity” means for our product, run experiments on new coding models and UX, and pinpoint where the model helps or hurts across languages and tasks. Your insights will directly shape how an entire industry builds software. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Codex product team to discover opportunities that improve developer outcomes and growth Design and interpret A/B tests and staged rollouts of new coding models and product features Define and operationalize metrics such as suggestion acceptance, edit distance, compile/test pass rates, task completion, latency, and session productivity Build dashboards and analyses that help the team self-serve answers to product questions (by language, framework, repo size, task type) Diagnose failure modes and partner with Research on targeted improvements (model quality signals, user feedback, evals) You might thrive in this role if you have 5+ years in a quantitative role at a developer-facing or high-growth product Fluency in SQL and Python; comfort with experiment design and causal inference Experience defining product metrics tied to user value Ability to communicate clearly with PM, Eng, and Design—and to influence product direction You could be an especially great fit if you have Strong programming background; ability to prototype, run simulations, and reason about code quality Familiarity with IDE/extensi
About the Team Our infrastructure team helps deliver OpenAI’s most capable models and products to the world by scaling infrastructure and turning demand into useful FLOPS. We collaborate across research, engineering, design, and business to turn cutting-edge AI advancements into impactful, real-world applications. Our team ensures the right compute is available—at the right time and place—to support some of the world’s most demanding workloads. We empower all of OpenAI’s products and research by scaling the infrastructure behind them. Our work makes it possible to launch new models and products reliably and at scale. About the Role As a Data Scientist on the Infra team, you will play a key role in shaping how we scale the infrastructure that powers OpenAI’s products and research. This is critical as we operate one of the largest and most advanced compute fleets in the world, supporting millions of users and businesses globally. We focus on aligning infrastructure measurement, planning, scaling, allocation, and efficiency to drive measurable impact across the company. You should expect to guide the definition of foundational datasets for infrastructure resources, develop metrics that inform key decisions, build forecasting and optimization models, and establish source of truth dashboards and analyses that enable teams to understand and improve infra usage. Most importantly, you should expect to be a core partner to engineering, research, and product teams in shaping the infrastructure that powers everything OpenAI builds. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain foundational datasets and metrics that reflect infrastructure usage, efficiency, and scaling. Develop forecasting and optimization models to support infra planning and resource allocation. Partner with engineering, research, and product teams to shape infrast
About the Team OpenAI’s Financial Engineering (FinEng) team powers how revenue flows through our products—pricing & packaging, checkout, payments, subscriptions, and the financial infrastructure behind them. We partner with Product, Engineering, Risk, Finance, and Go-to-Market to make paying for OpenAI products seamless, reliable, and efficient worldwide. About the Role As a Data Scientist on FinEng, you’ll own the analytics and experimentation that improve our checkout and payments , subscriptions , and pricing & monetization systems. You’ll define the metrics that matter, build the source-of-truth data assets, and design experiments that increase conversion, reduce churn and payment failures, and expand global payment method coverage. Your work will directly influence revenue, customer experience, and how we scale internationally. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Own checkout & payments analytics and experimentation across methods and locales (e.g., bank transfers, emerging rails), improving conversion while monitoring risk and latency. Build and run the experimentation program for in-house checkout—define success metrics and guardrails, execute staged rollouts, and use offline incrementality when online tests aren’t feasible. Create operational visibility and source-of-truth data with FinEng Data Engineering—land team-level metrics, SLAs, and self-serve dashboards that drive proactive action. Lead subscription, retention, and monetization analytics—ship launch-readiness for new subscription features, reduce involuntary churn (e.g., targeted retrials/nudges), and develop elasticity/FX frameworks toward pricing optimality. You might thrive in this role if you have 5+ years in a quantitative role (data science, product analytics, or experimentation) in high-growth or fintech environments Fluency in SQL and Python ,
About the Team OpenAI’s Platform team powers how millions of developers and enterprises build with our models. We provide APIs and agentic solutions used by global startups and fortune 500s. We work closely with product, engineering, design, and go-to-market to build a world-class platform that pushes the frontier of AI capabilities. About the Role As a Data Scientist on the Platform team, you will drive a data-driven culture for OpenAI’s API and B2B solutions. You’ll define the metrics that matter for developer success and enterprise value, measure the impact of new models and features, and partner with PMs and engineers to improve model quality, reliability, latency, and cost. Your work will shape how thousands of products adopt agentic AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Platform product team as a trusted partner, uncovering ways to improve developer experience, reliability, and usage growth Define north-star metrics across the developer funnel (activation, retention, growth), as well as latency/cost guardrails for new features and models Design and interpret A/B tests and controlled rollouts (e.g., new model versions, pricing/limits, new API features, new B2B products) Build source-of-truth dashboards and self-serve data tools for product, engineering, and go-to-market teams Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality) You might thrive in this role if you have 5+ years in a quantitative role in ambiguous, high-growth environments (platforms, APIs, or B2B products a plus) Depth in SQL and Python, with a track record proposing, designing, and running rigorous experiments Experience defining and operationalizing metrics from scratch (including reliability/latency/cost and safety) Strong cross-functional communication with PMs, enginee
About the Team OpenAI’s mission is to ensure that general-purpose artificial intelligence benefits all of humanity. We believe that achieving our goal requires real world deployment and iteratively updating based on what we learn. The Protection Scientist Engineer, Integrity team supports this by identifying and investigating misuses of our products – especially new types of abuse. This enables our partner teams to develop data-backed product policies and build scaled safety mitigations. Precisely understanding abuse allows us to safely enable users to build useful things with our products. About the Role Protection Science Engineering is an interdisciplinary role mixing data science, machine learning, investigation, and policy/protocol development. As a Protection Scientist Engineer within Integrity and Investigations, you will be responsible for designing and building systems to proactively identify and enforce on abuse on OpenAI’s products. This includes ensuring we have robust abuse monitoring in place for new products, sustaining monitoring for existing products, and prototyping and incubating systems of defense against our highest risk harms. You will also respond to and investigate critical escalations, especially those that are not caught by our existing safety systems. This will require expert understanding of our products and data, and involves working cross-functionally with product, policy, and engineering teams. This role can be based in either our San Francisco, or NY office and includes participation in an on-call rotation that will involve resolving urgent escalations outside of normal work hours. Some investigations may involve sensitive content, including sexual, violent, or otherwise-disturbing material. In this role, you will: Scope and implement abuse monitoring requirements for new product launches. Improve processes to sustain monitoring operations for existing products, including developing approaches to automate monitoring subtasks. Prototyp
Job Description: Data Scientist, B2B Demand Generation, Growth & Measurement About the Role We are hiring a Data Scientist to lead measurement, experimentation, and decision science for B2B marketing demand generation. You will help us understand which marketing investments create incremental demand, qualified pipeline, and revenue and how to scale them efficiently. Our mandate is to build a rigorous, full-funnel view of how B2B marketing creates demand and moves prospects from awareness and engagement to qualified opportunities, closed-won revenue, and expansion. You will shape how we measure marketing impact and influence across channels, campaigns, audiences, and account segments. In this role, you will partner closely with B2B Marketing, Demand Generation, Growth, Sales, RevOps, Finance to connect marketing activity to qualified pipeline, customer acquisition, and efficient revenue growth. What You’ll Do Define north-star, leading, and guardrail metrics for B2B demand generation, including account engagement, qualified leads and opportunities, sourced and influenced pipeline, conversion rates, pipeline velocity, and incremental ARR. Design and execute measurement and experimentation strategies across channels and campaigns, using randomized tests, audience or geographic holdouts, lift studies, quasi-experimental methods, and other causal approaches suited to long B2B sales cycles. Analyze channel, audience, campaign, creative, content, landing-page, and account-segment performance to identify the drivers of qualified demand, funnel conversion, pipeline quality, and incremental revenue. Partner with Marketing, Sales, RevOps, Finance, Product, and Engineering to improve instrumentation, campaign taxonomy, CRM data quality, lead-to-account matching, and the operating cadence for acting on measurement insights. Build AI-native measurement and decision-support workflows, using LLMs and agents to synthesize campaign performance, surface growth opportunities, and h
About the Team The Preparedness team is an important part of the Safety Systems org at OpenAI, and is guided by OpenAI’s Preparedness Framework . Frontier AI models have the potential to benefit all of humanity, but also pose increasingly severe risks. To ensure that AI promotes positive change, the Preparedness team helps us prepare for the development of increasingly capable frontier AI models. This team is tasked with identifying, tracking, and preparing for catastrophic risks related to frontier AI models. The mission of the Preparedness team is to: Closely monitor and predict the evolving capabilities of frontier AI systems, with an eye towards misuse risks whose impact could be catastrophic to our society Ensure we have concrete procedures, infrastructure and partnerships to mitigate these risks and to safely handle the development of powerful AI systems Preparedness tightly connects capability assessment, evaluations, and internal red teaming, and mitigations for frontier models, as well as overall coordination on AGI preparedness. This is fast paced, exciting work that has far reaching importance for the company and for society. About the Role We’re hiring a Data Scientist to help build, evaluate, and continuously improve mitigations that prevent extreme harms from AI systems. This role is for an experienced, highly autonomous individual contributor who can take ambiguous problem statements, structure rigorous analyses, and translate findings into actionable product and policy changes. This position goes beyond “running evals.” You’ll help create mitigation intelligence and monitoring systems that enable OpenAI to detect issues early, measure effectiveness over time, and reduce both over-blocking (unnecessary friction) and under-blocking (missed harm). What You’ll Do Evaluate and improve mitigation systems, including classifiers and detection pipelines across domains (e.g., biosecurity, cybersecurity, and emerging risk areas). Diagnose false positives and fa
About the Team The Future of Computing Research team is an applied research team within the Consumer Devices group focused on developing new methods, models, and evaluation frameworks that support our vision for the future of computing. We work at the frontier of multimodal AI, helping turn emerging model capabilities into product experiences that are useful, delightful, and worthy of long-term trust. Our work explores a new class of AI systems that can learn over time, adapt to individuals, and support people in the flow of daily life. This includes long-term memory, user modeling, and personalization systems that are aligned not just with immediate satisfaction, but with a person’s broader goals, values, and well-being. We work closely across research, engineering, design, product, and safety to define what it means to build AI systems that know you over time, act at the right moment, and help in ways that are context-aware, respectful, and demonstrably beneficial. About the Role We are looking for a Research Engineer / Scientist to join the Future of Computing Research team to work on RLHF and post-training for personalized, multimodal AI systems. This role will focus on building the learning and evaluation foundations that help models become more context-aware, adaptive, and useful over time. You will work on problems such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that must make high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded: success is not just higher benchmark performance, but better model behavior in real-world use. The ideal candidate is excited about pushing beyond one-turn assistant behavior toward systems that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value. Internally, that maps closely to the need for careful reward design, feedback loops, and evaluation frameworks that test whether i
🚀 About WRITER WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI. Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI. 📐 About the role AI research at WRITER isn't just about publishing papers — it's about building the scientific foundation that powers some of the most ambitious enterprise AI deployments in the world. As an AI research scientist, you'll be at the center of that work. You'll drive a high-impact research agenda focused on large language models, agentic reasoning, and the system-level capabilities that make AI genuinely useful at enterprise scale. This is a rare opportunity to do research that matters twice over — advancing the field and shipping directly into products used by hundreds of thousands of people every day. We're at an inflection point. Enterprises are moving from experimenting with AI to deeply embedding it across their operations, and WRITER's models are the engine making that possible. The work you do here — on post-training, planning, multi-step reasoning, and agentic workflows — will directly shape how the next generation of enterprise AI behaves, performs, and scales. You
$170K – $225K/yr
About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! Prior to applying please note: W e are currently unable to provide visa sponsorship for this position (including H-1B, OPT, F1, CPT or other employment-based visas). Candidates must be legally authorized to work in the United States without employer sponsorship now or in the future. This role is hybrid requiring 2 days in office at our San Francisco hub every Tuesday & Wednesday (located at 130 Sutter St, San Francisco, CA). About the Role Data Science plays a crucial role in driving impact at Taskrabbit. We are seeking a highly skilled and motivated Staff Data Scientist to join us, working closely with cross-functional teams from Product teams and occasionally Commercial Operations to provide data-driven insights and solutions that enhance our products and accelerate growth while minimizing marketplace losses. What you will work on Be a strategic thought partner with stakeholders in Product and occasionally Commercia
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. About the Team At Embedded Insights, we find the best machine learning opportunities for external products and internal systems, and collaborate with cross-functional partners to bring them to life. We are a central team of Machine Learning Engineers and Data Scientists. We embed with partner teams to build and apply machine learning models that improve internal decision-making and power the Plaid product suite. About the Role You will be the first Data Scientist on the Embedded Insights team, part of Plaid’s Data organization. You will establish the analytics and metrics backbone for a team supporting a diverse set of internal and external products. You will help drive better decision-making, support machine learning model development, and contribute directly to the health of the Plaid network and the quality of Plaid’s products. Your day-to-day work will include: Analyzing entities across the Plaid network to understand behavior and identify opportunities, anomalies, and risks. Creating foundational metrics, dashboards, and monitoring systems that provide a clear view of network health and machine learning model performance. Evaluating the value and performance of machine learni
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption
About the Team OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make rigorous, evidence-based talent decisions and ensures that the systems supporting those decisions are valid, reliable, fair, and accountable. About the Role As a People Data Scientist focused on AI fairness and bias testing, you will help establish how OpenAI evaluates AI-assisted People systems and high-impact talent processes. You will design and conduct rigorous assessments to identify, measure, and mitigate potential bias across the lifecycle of models, agents, decision-support tools, and automated workflows. Your work will span the entire employee life-cycle, such as hiring, performance, promotion, employee development, workforce planning, etc. You will evaluate both technical systems and the broader human-AI decision processes in which they operate, examining not only model performance but also data quality, measurement validity, differential outcomes, human oversight, and unintended consequences. We’re looking for an experienced data scientist or applied researcher who can translate complex fairness questions into defensible evaluation strategies, scalable testing infrastructure, and clear recommendations for technical teams and senior leaders. This role is preferred to be based in San Francisco, CA. In this role, you will: Define and lead fairness and bias-testing strategies for AI-assisted People processes, models, agents, and decision-support systems from development through deployment and ongoing monitoring. Design rigorous algorithmic audits and validation studies, including adverse-impact analysis, subgroup and intersectional evaluation, error-rate analysis, calibration, measurement invariance, reliability, criterion-related validity, and sensitivity testing. Identify the appropriate fairness criteria for each use case, evaluate tradeoffs among competing definitions
About the Team OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make better, evidence-based talent decisions. About the Role As a People Research Scientist, you will bring deep expertise in research design, measurement, experimentation, and applied data science to OpenAI’s most important People programs. You will design studies, evaluate people processes, and help leaders better empower employees, strengthen organizational systems, and deliver exceptional employee experiences. This is a high-ownership individual contributor role combining hands-on research, methodological leadership, and scalable people science capabilities. We’re looking for an experienced researcher who can turn ambiguous People questions into rigorous designs, validated insights, and actionable recommendations. This role is based in San Francisco, CA or Mountain View, CA, with occasional travel to our San Francisco office. What You’ll Do: Design rigorous research and evaluation strategies for recruiting, organizational health, manager effectiveness, employee experience, and talent outcomes. Apply advanced statistical modeling, machine learning, and research methods to inform program design, evaluate effectiveness, and quantify business impact. Partner with People Operations, data engineering, and people systems teams to define data requirements, improve data quality, establish documentation standards, and ensure research datasets are governed, reproducible, and privacy-preserving. Build scalable people science infrastructure, including self-service agentic tools, automated validation workflows, reusable research datasets and analytical pipelines. Develop research playbooks that establish rigorous standards for study design, measurement, validation, and documentation, enabling high-quality, repeatable, and scalable research across the organization. Communicate findings through c
Other cities to consider
More places hiring for this role
Get new scientist jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime