Jobs in United States

Data Client Onboarding Specialist in San Francisco

581 active opportunities · Updated October 2026

Explore current data client onboarding specialist jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,

PythonKubernetesGitRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Search research team focuses on building the systems that help AI systems find, retrieve, and use information from the world. We aim to make answers more useful and grounded for more than a billion ChatGPT users. About the Role We’re looking for a Technical Program Manager to lead a broad portfolio of research and engineering programs that power search. You’ll partner closely with researchers, engineers, and product leaders to turn ambitious goals into clear plans, resolve dependencies, and move complex technical work forward. This role combines technical depth, product judgment, and hands-on execution. You’ll work across retrieval, indexing, and model improvements, while collaborating with policy, legal, and external data partners. You’ll help teams make informed tradeoffs and build practical ways of working that support a fast-moving research environment. This role is based in San Francisco, CA. In this role, you will: Lead programs across model training, retrieval, large-scale indexing, and search infrastructure. Translate evolving goals into prioritized workstreams with clear owners, milestones, dependencies, and resource needs. Partner with research, engineering, and product leads to define requirements and make tradeoffs across scope, quality, performance, timelines, and cost. Establish program success metrics and use them to guide priorities and track improvements in coverage, answer quality, responsiveness, and trust. Identify technical and cross-functional risks early, drive blockers to resolution, and communicate progress and decisions clearly to teams and leadership. Coordinate with product, policy, legal, and external partners on data access, use, and presentation, helping teams resolve decisions that span technical and non-technical domains. Manage dependencies with data providers and build repeatable processes that help research and engineering teams execute effectively as the search effort grows. You might thrive in this role if you

AWSRestAIGo
M
📍 San Francisco, United States· Full-time
✓ High-confidence listingCompany trend -55.6%
Quick readStrong listing-quality and freshness signals

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the products they build. By combining powerful analytics with AI that knows your business, Mixpanel helps teams see what’s working, diagnose what’s not, and decide what to build next. Learn more at mixpanel.com . About Mixpanel Mixpanel turns data clarity into innovation. Trusted by more than 29,000 companies, including Workday, Pinterest, LG, and Rakuten Viber, Mixpanel’s AI-first digital analytics help teams accelerate adoption, improve retention, and ship with confidence. Powering this is an industry-leading platform that combines product and web analytics, session replay, experimentation, feature flags, and metric trees. Mixpanel delivers insights that customers trust. Visit mixpanel.com to learn more. About The Team Mixpanel Engineering is a small, fast-moving team focused on delivering real value to customers. We build powerful AI-powered product analytics while obsessing over clarity, simplicity, and delight. Engineers here own problems end to end. You can move across the stack to ship impact without being blocked by silos or heavy process. Product innovation drives our business, and product engineering teams own that responsibility. Our OLAP engine queries over 500 trillion events; a typical blob storage system we interact with processes 300 PiB/month at 1.2 Tbps sustained, and we run many of them across the world. The Data Runtime team owns the data execution layer that powers every Mixpanel product. We ensure that every customer query runs fast, cheap, and reliably, at any scale. This is an exciting time to join. Mixpanel's agentic and AI-first products are driving rapid growth in query volume, and Data Runtime is making the big bets that power it. We’re investing in elastic query compute and a distributed file cache that will let us scale query workloads dramatically without scaling cost with them. We

PythonSQLAWSAzure
P
📍 San Francisco, CA, United States· Remote
✓ Quality checkedCompany trend -85.6%

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . About tvScientific tvScientific is the first and only CTV advertising platform purpose-built for performance marketers. We leverage massive data and cutting-edge science to automate and optimize TV advertising to drive business outcomes. Our solution combines media buying, optimization, measurement, and attribution in one, efficient platform. Our platform is built by industry leaders with a long history in programmatic advertising, digital media, and ad verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business. As a Data Engineer at tvScientific, you will be a key player in implementing the robust data infrastructure to power our data-heavy company. You will collaborate with our cross-functional teams to evolve our core data pipelines, design for efficiency as we scale, and store data

SQLAWSMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure: Owner Furnished Equipment to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastr

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to support the next generation of frontier AI systems. Through a combination of strategic partnerships and self-built data center campuses, we are scaling the power, cooling, electrical, mechanical, and controls infrastructure needed to deliver compute at unprecedented scale. The Commissioning organization is responsible for ensuring this infrastructure is safely tested, validated, integrated, and transitioned into reliable operations. For our self-build campuses, the team operates through a hybrid delivery model: OpenAI provides commissioning leadership, discipline ownership, governance, and project integration, while commissioning partners provide field and test engineering capacity to support inspections, startup, testing, and turnover. About the Role We are seeking a Commissioning Project Lead to own the commissioning strategy and execution for a large-scale, self-build data center project. You will lead the overall commissioning program from early construction planning through startup, functional testing, integrated systems testing, and final turnover. You will establish the commissioning execution plan, integrate commissioning activities into the master project schedule, coordinate multidisciplinary readiness, and lead the vendor commissioning partners providing field and test engineering capacity. This role serves as the primary commissioning interface to project leadership, construction management, contractors, equipment vendors, operations, and commissioning partners. You will be responsible for creating clarity across organizations, identifying readiness and schedule risks early, and ensuring the facility progresses through testing and turnover against clearly defined acceptance criteria. The role will initially support planning and coordination in a hybrid capacity and transition to full-time onsite presence as construction, inspections, startup, testing, and t

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Applied team brings OpenAI’s technology to the world through products used by hundreds of millions of people and by developers and businesses building on our APIs. We work across research, engineering, product, policy, safety, and operations to deploy frontier AI systems responsibly and safely. The Trust & Safety Data Engineering team builds the data foundations that help OpenAI understand, detect, investigate, and mitigate abuse and safety risks across our products. We partner with Integrity, Investigations, Safety Systems, Product Policy, Privacy, Data Science, Engineering, and Data Platform to create reliable, privacy-safe datasets and pipelines for fraud and abuse detection, enforcement workflows, safety measurement, ML feature generation, launch readiness, and transparency reporting. About the Role We are hiring a Technical Lead Manager to lead and grow the Trust & Safety Data Engineering team. This is a hands-on leadership role for someone who can set strategy, shape data architecture, align senior stakeholders, coach engineers, and drive execution on high-impact data systems. You will help turn fragmented launch and incident support into durable, reusable, privacy-safe data foundations that Trust & Safety teams can rely on. The systems your team builds will help OpenAI detect risk, investigate abuse, power operational workflows, develop and evaluate safety models, measure interventions, support product launches, and report accurately on platform integrity. In This Role, You Will Lead and grow a high-performing Trust & Safety Data Engineering team. Define the roadmap and technical strategy for Trust & Safety data systems. Build canonical, privacy-safe datasets and pipelines for abuse detection, fraud detection, risk signals, enforcement, scaled review, transparency reporting, and safety monitoring. Create reusable foundations for Trust & Safety model development, including features, labels, training data, backtesting,

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Software Engineer, Distributed Data Systems, you will design, build, and operate some of the largest distributed data systems in the world. You will be responsible for the end-to-end stack to deliver and consume top-quality data for robotics training at exabyte-scale. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI’s rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable, large-scale systems in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure such as exabyte-scale distributed data processing, data selection, automated labeling, and training data loaders. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. Deliver the best possible data for training robotics models. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented a

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. About the Role We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. Design and iterate on user-facing tools and backend services to support high-quality data workflows Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal r

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside leading cloud providers, engineering firms, construction partners, utilities, and equipment manufacturers, we are delivering hyperscale AI campuses that enable the next generation of frontier AI models. The Strategic Sourcing team develops and executes the commercial strategies that ensure our infrastructure programs have reliable access to the equipment, materials, and strategic partners needed to deliver at unprecedented scale. We partner closely with Infrastructure Delivery, Capacity Planning, Design Engineering, Hardware Operations, Finance, Legal, and our external suppliers to build a resilient global supply network capable of supporting Industrial Compute's long-term growth. As we continue expanding globally, strategic sourcing becomes a critical competitive advantage, ensuring our infrastructure programs remain cost-effective, resilient, and capable of executing against aggressive deployment timelines. About the Role We are seeking a Strategic Sourcing Manager, Data Center Infrastructure to lead sourcing strategy for the critical infrastructure systems that power Industrial Compute campuses. This role will develop commercial strategies, negotiate strategic supplier agreements, and manage relationships across engineering, construction, manufacturing, and infrastructure partners responsible for delivering mission-critical facilities. You will work closely with Infrastructure Delivery, Capacity Planning, Engineering, Finance, Construction, and external suppliers to ensure Industrial Compute has the capacity, supplier relationships, and commercial frameworks required to support rapid global expansion. The ideal candidate has experience sourcing major infrastructure systems for hyperscale data centers, mission-critical facilities, industrial construction, semiconductor manufacturing, energy infrastructure, or similarly comple

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale. Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets. By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life. About the Role We are looking for an engineer to design and implement the dataset infrastructure that powers OpenAI’s next-generation training stack. You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks. In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. In this role, you will: Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory. Build proactive testing and scale validation pipelines for dataset loading at GPU scale. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience. Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt. Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized. Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training). Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets. You might thrive in this role if you: Have strong engineering fundamentals with experience in distributed systems, data pipelines, or infrastructure.

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines, collaborate closely with researchers to translate requirements into robust systems, and harden pipelines that serve as the backbone for OpenAI's rapid iteration cycles. We’re looking for engineers who are detail-oriented, have strong experience with distributed systems, and excel at building reliable infrastructure in high-stakes environments. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security. Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient. Partner with researchers to deeply understand requirements and translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation. You might thrive in this role if you: Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data. Are detail-oriented and bring rigor to building and maintaining reliable systems. Demonstrate excellent software enginee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team At OpenAI, Trust & Safety Operations is central to protecting OpenAI’s platform, customers, and the public from abuse. We partner closely with Product, Engineering, Legal, Policy and Go To Market teams to identify emerging risks, build and mature enforcement systems, and ensure high-integrity operations while delivering a great user experience at scale. We’re building the Monetization Trust & Safety Operations team to ensure OpenAI can grow advertising in a way that is safe, trusted, and sustainable—for users, advertisers, and the business. This team sits at the intersection of operational scale, product risk, and rapid revenue growth, designing systems and operations that enable ads to scale without compromising user trust or safety. It’s critical to us that our Ads product be built in a way that corresponds to our Ads principles , and this team is key to that. About the Role We’re looking for a senior operator with strong analytical instincts to help build and scale Monetization Trust & Safety Operations at OpenAI. In this role, you’ll flex across the team’s highest-priority data and operational needs—from reporting and dashboard insights to budget and capacity planning, project-based analysis, and data automation —while partnering closely with Product, Policy, Engineering, Legal, Go To Market, and Data Science and Data Engineering teams. This role sits at the intersection of strategy, execution, and data: you’ll define ambiguous problems, query and validate data, build decision-support systems, and translate operational signals into clear recommendations and scalable, AI-first solutions. You should be comfortable moving from a high-level question to a rigorous analysis, a useful dashboard, an automated workflow, or a durable operating mechanism. As OpenAI introduces new revenue-generating formats and partnerships, you’ll help the team understand where risks, capacity constraints, quality gaps, and opportunities are emerging. You’ll brin

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late

PythonAWSRestMachine Learning
🔔

Get new data client onboarding specialist jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime