About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize
Jobs in United States
Data Operations Lead in San Francisco
581 active opportunities · Updated October 2026
Showing
15 jobs
Explore current data operations lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Baseten is seeking talented and experienced Software Engineers to join our Observability team within the Infrastructure organization. As an early member of the Observability Team, you will be pivotal in building and shaping the observability experience for our internal and external customers. By joining this team, you’ll have a direct impact on the reliability and operational excellence of Basetens product systems. As Baseten scales its infrastructure across different cloud providers and diverse hardware, the volume and complexity of operational data is growing by orders of magnitude. This team is responsible for building high-throughput ingest pipelines, cost-efficient storage, and agentic diagnostic tools to ensure that we can detect, diagnose, and resolve issues in minutes rather than hours, even as the systems they operate become more complex. RESPONSIBILITIES Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across Baseten’s multi-cloud infrastructure Own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth Build instrumentation libraries, SDKs, and integrations that make it easy for engineering teams to emit high-quality telemetry from their services Drive alerting and SLO infrastructure that enables teams to define, monitor, and respond to reliabi
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re hiring a Data Scientist to help build and scale our internal analytics capabilities. This is a foundational role where you’ll create dashboards, data models and insights to power business and product teams alike. You’ll collect requirements, define key metrics, and deliver insights directly to stakeholders. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape Baseten’s product and strategy. RESPONSIBILITIES Build and maintain production-grade dbt models and dashboards across multiple functions with a focus on accuracy, simplicity and user experience. Define and instrument core metrics around ROI, product adoption, customer lifecycle, capacity, availability, revenue and costs. Ingest and transform raw data using tools like dbt, Airbyte, and BigQuery. Partner with Engineering, Finance, Marketing, and Sales teams to understand goals and translate them into data solutions REQUIREMENTS 5+ years of experience in analytics engineering, data analysis, analytics, data science or a related role Advanced SQL and dbt skills, with a record of building models, tests, semantic layers and lineage in a cloud data warehouse. Prior experience supporting complex cross-functional projects across GTM, Finance and Engineering across various stages of the customer journey. Experience building dashboards and self-serve analy
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Product Data Scientist to establish how product decisions at Baseten are made with data. You'll work directly with Product and Engineering, alongside GTM to determine measurement, strategy, experimentation and implementation. This is a foundational, hands-on role. You'll define what success looks like across a technical, usage-based platform and turn ambiguous questions into analyses, forecasts, and experiments that shape product strategy. You'll work from clickstream and product events through inference telemetry and observability data, helping Baseten make faster decisions about reliability, performance, adoption and developer experience. RESPONSIBILITIES Partner directly with Product and Engineering: frame the questions that matter, define success criteria, and turn analysis into roadmap, launch, and prioritization decisions. Define how product success is measured: establish metrics across activation, adoption, retention, expansion, reliability and user experience. Support experimentation and launches: design measurement plans, analyze A/B experiments and controlled rollouts, and translate results into product decisions. Diagnose reliability and scaling behavior: join customer signals with request, replica, deployment, and cluster telemetry to find patterns in release bottlenecks, unhealthy replicas, and models without traffic. Define the enterprise customer journey and measure feature adoption
About the Team The GTM Intelligence Solutions team builds the data and decision systems that help customer-facing teams take the right action at the right time. We combine product telemetry, commercial data, customer context, and field activity to identify account health and opportunity, recommend actions and use cases, deliver intelligence through field-facing products and agent workflows, and measure what happens next. We’re looking for a Data Scientist to help build the next generation of GTM intelligence at OpenAI. You will own a flexible portfolio of high-impact decision data products and work closely with Technical Success and other GTM teams to ensure the work drives better decisions. About the Role As a Data Scientist on GTM Intelligence Solutions, you will define and build the intelligence systems that help customer-facing teams prioritize accounts, identify risks and opportunities, choose interventions, and understand what worked. You will set the roadmap and methodology, build canonical features, ship reliable production workflows, monitor quality and adoption, and improve the systems using field feedback and business outcomes. This role combines hands-on technical depth with strong product and business judgment. You should be as comfortable writing production Python and advanced SQL, defining durable data contracts, and operating decision products as you are evaluating a ranking approach or designing an experiment. You will personally ship reliable first versions and partner with Analytics Engineering and Data Engineering when work requires shared infrastructure or additional scale. In This Role, You Will Set the roadmap and methodology for GTM intelligence and decision products, using deep stakeholder discovery to probe beyond stated requests, uncover the underlying decisions, workflows, constraints, and measures of success, and translate them into measurable systems. Own the full lifecycle of intelligence products, including feature definition, methodo
About the Team The Consumer Devices team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. About the role We are looking for an Operating Systems Engineer to build and harden the OS foundations for OpenAI products. We are especially interested in experienced, passionate, and innovative operating systems developers who thrive on building foundational platform software and solving hard problems in security, privacy, performance, power, and reliability. You will work across the OS kernel, core OS services, security and privacy primitives, performance and power, and the frameworks that connect applications and UI to the system. This role emphasizes deep debugging and systems ownership from development through production. You will collaborate closely with embedded, firmware, hardware, application, and product engineering teams. Experience with hardware bring-up is a plus, but not required. What you will do Work on end-to-end OS capabilities spanning the OS kernel, userspace services, application frameworks, UI toolkits, and application-facing APIs. Develop, integrate, and maintain OS components, both kernel-bound and in userspace, including scheduling, memory management, filesystems, drivers, IPC/RPC mechanisms, and security-relevant subsystems. Build and maintain core OS services and daemons (init, service management, device discovery, networking primitives, time, logging, update hooks, crash handling, and so on). Design and implement security and privacy mechanisms: Secure boot and measured boot integration points (where applicable). Mandatory access control and sandboxing. Secrets management, secure storage, key handling, and least-privilege service design. Privacy-preserving telemetry, data minimization, and user-consent oriented system behaviors. Establish a perfo
What you’ll do Design and implement secure cloud pipelines that ingest very large scan datasets (multi-terabyte), reliably and resumably. Build orchestration for GPU-accelerated reconstruction and analysis with strong retry semantics, idempotency, and cost controls. Define end-to-end data lifecycle for medical imaging: raw vs intermediate vs derived artifacts, retention policies, and reproducibility. Implement security + compliance primitives appropriate for HIPAA/PHI: encryption in transit/at rest, key management, least privilege, audit logs, and access reviews. Build operational tooling: monitoring, alerting, runbooks, and incident-driven improvements for a growing device fleet. What we’re looking for Strong experience with cloud batch/queueing/orchestration, storage systems, and data pipeline reliability. Experience shipping production systems that handle large data volumes and failure-prone networks. Practical security mindset (least privilege, secrets, audit logging) and comfort operating in compliance-constrained environments. Useful experience Building reliable data pipelines at scale (queues/orchestration, resumable uploads, GPU batch execution) with strong observability. Security + privacy by default: encryption, least-privilege access, auditing, and practical HIPAA/PHI guardrails. Owning the “boring” backend details that keep a lean team moving: schemas/migrations, cost controls, retries, and runbooks. Understanding compute tradeoffs across hardware options, and specifying appropriate cloud resources.
About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Operating Systems team is critical in this mission, turning sophisticated hardware and AI capabilities into a reliable, trusted platform. Security is central to that work: we define trust boundaries, integrate hardware-backed protections, isolate sensitive context, and establish the guardrails that let AI applications and agents act safely, privately, and under user control. About the Role We’re looking for a Software Security Architect to define the security architecture for OpenAI’s next-generation operating system. You’ll work alongside hardware security architects and partner with operating system, silicon, firmware, privacy, and product teams to protect users, their devices, and their data. This is a senior, hands-on role for someone who can connect operating system internals, hardware-backed security, and real-world product constraints. Your work will shape the platform’s trust model, protect sensitive information, and set the technical foundations for AI that is safe, private, and under user control. In this role, you will: Define the operating system’s security architecture, trust boundaries, privilege model, and protections for sensitive user data. Partner with hardware security architects to integrate roots of trust, secure elements, trusted execution environments, and processor security capabilities into the operating system. Desig
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The Security Governance, Risk, and Compliance (GRC) team is part of Plaid’s security organization, focused on enabling the business by proactively managing information security risks and maintaining effective controls. Our mission is to reduce the likelihood and impact of security risks while operating a robust assurance program that builds trust with our customers, consumers, and data partners. We own Plaid’s security compliance frameworks, run our audits and risk programs, and partner across the company to keep Plaid’s platform secure, resilient, and aligned with industry and regulatory expectations. GRC Engineering is how we make all of that scale — turning compliance into code, evidence into telemetry, and audits into a continuous, automated capability. The Role: You will own GRC Engineering at Plaid — a foundational, high-ownership role defining an emerging discipline from the ground up. Today most of our compliance work is manual and point-in-time; you will turn it into an engineered system that is continuous, data-driven, and scalable, and set the technical direction for the field. You will: Define the discipline and the architecture — how GRC Engineering works at Plaid, not just execute with
About the Team Enterprise Verticals builds role-specific ChatGPT Work experiences for high-value enterprise workflows. We combine product engineering, plugins and skills, connectors, data, evaluations, and customer evidence to turn useful demos into reliable daily work. This opening sits within the Technology vertical inside Enterprise Verticals. The group focuses on repeatable workflows for people at technology companies, beginning with functions such as data and analytics, sales, and design, and carries the shared platform needs—tool integration, permissions, quality measurement, and safe rollout—across those experiences. We work closely with Design, Research, GTM, Security, and platform teams, as well as with customers and design partners. Success means that people can reach a trustworthy first result, understand what the system did, and keep using the workflow—not merely that a prototype exists. About the Role We are looking for a full-stack product engineer who can own ambiguous enterprise workflows end to end: understand a customer problem, shape the product, build across frontend and backend, work through platform dependencies, instrument quality, and learn quickly with design partners. You will build across ChatGPT Work surfaces, services, plugins, connectors, and data or permission boundaries when the experience requires it. You will make quality and rollout observable through evaluations, product and operational signals, and clear fallback or rollback paths. This is a product-engineering role for someone who can move between user problems and system details without losing ownership of either. The strongest candidates will be able to turn specific customer evidence into a generalizable product, explain scope and architecture tradeoffs, and carry a feature from an early prototype through a bounded production rollout. In this role, you will: Build and ship role-specific workflows across ChatGPT Work surfaces, services, plugins, and connectors. Turn customer a
We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We're hiring a Marketing Analytics Manager to establish how Marketing at Baseten makes decisions with data. Marketing at Baseten is scaling fast: more spend, more campaigns, more model launches and more inbound. This is a foundational, hands-on role and the first dedicated Marketing Analytics hire. You'll work directly with Demand Gen, Field Marketing, MOps and Product Marketing alongside GTM, Finance, Product and Engineering to stitch together activities and outcomes across the funnel. You'll build the data models that connect acquisition, engagement, activation, product usage and revenue. You’ll design dashboards, tools, semantic layers and plugins that enable teams to answer questions. Along the way, you’ll develop an understanding of how customers use Baseten and with this, define how our systems and product can improve. RESPONSIBILITIES Define how Marketing success is measured: establish metrics across audience growth, acquisition, activation, engagement, usage, pipeline and revenue. Build the marketing data foundation: ingest and model data across Salesforce, HubSpot, Google Analytics, advertising platforms, web, email, product and third-party sources. Create a medallion architecture that connects the prospect and customer journeys across systems, from first touch through signup, onboarding, activation and expansion. Understand our audiences: develop audience and segmentation frameworks based on customer a
About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a researcher focused on our embedding retrieval efforts. You’ll work with a a team of world-class research scientists and engineers developing foundational technology that enables models to retrieve and condition on the right information, at the right time. This includes designing new embedding training objectives, scalable vector store architectures, and dynamic indexing methods. This work will support retrieval across many OpenAI products and internal research efforts, with opportunities for scientific publication and deep technical impact. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Responsibilities Tackle embedding models and retrieval systems optimized for grounding, relevance, and adaptive reasoning. Collaborate with a team of researchers and engineers building end-to-end infrastructure for training, evaluati
About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a Software Engineer focused on building and scaling retrieval systems. You’ll work with a team of researchers and engineers to develop infrastructure that enables models to retrieve and act on the right information at the right time. This includes designing and operating indexing systems, retrieval pipelines, and serving layers. This work supports retrieval across OpenAI products and research, with direct impact on system performance, reliability, and scale. Responsibilities Build and scale retrieval infrastructure across indexing, serving, and query execution. Develop low-latency, high-throughput systems for real-time model interaction. Partner with research to productionize embedding and retrieval techniques. Support dense, sparse, and hybrid retrieval pipelines. Own system performance, reliability, and observability at scale. Collaborate across Pretraining, Inference, and Product teams to integrate retrieval end-to-e
About the Team The Intelligence & Investigations Engineering team builds systems that detect, analyze, and disrupt abuse across OpenAI’s products. We partner closely with the Child Safety team and cross-functional groups to protect users while advancing OpenAI’s goal of developing AI that benefits everyone. About the Role As a Fullstack Engineer focused on child safety, you’ll build data-intensive, AI-powered applications and infrastructure that enable operators and investigators to work effectively and responsibly. You’ll adapt quickly in ambiguous, fast-moving environments to deliver well-crafted, reliable tooling for high-severity safety work. *Candidates should understand this role involves exposure to sensitive and egregious content. In this role, you will: 4+ years of experience as a software engineer Prototype, build, and maintain intelligence systems that detect, triage, and enable efficient human review of possible high severity harm Work hand in hand with operators and investigators, designing and delivering systems that enable them to do their work faster, more accurately, and more safely. Develop across the stack: UIs, services, pipelines, and anything else required to solve the problems we face. Interact with partners across Product Policy, Platform Integrity, Safety Systems, and Research Contribute to the team’s technical strategy, especially for child safety related tools and systems Report on impact in a data-driven fashion You might thrive in this role if you: Have a strong software engineering foundation and enjoy owning systems end-to-end—from infrastructure and data ingestion to frontend tooling Are energized by working at the frontier of AI capabilities, integrating new models and APIs into practical systems Have experience building and operating large-scale data pipelines or search/retrieval systems Are proficient in Python and/or TypeScript, and familiar with tools like Spark, Kafka, Flink, data warehouses, and SQL Take a product-minded ap
Other cities to consider
More places hiring for this role
Get new data operations lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime