Jobs in United States

Data Entry Specialist in United States

2,501 active opportunities · Updated October 2026

Explore current data entry specialist jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Software Engineer to join our Data Acquisition team. Responsibilities: Own and lead engineering projects in the area of data acquisition including web crawling, data ingestion, and search. Collaborate with other sub-teams, such as Data Processing, Architecture, and Scaling, to ensure smooth data flow and system operability. Work closely with the legal team to handle any compliance or data privacy-related matters. Develop and deploy highly scalable distributed systems capable of handling petabytes of data. Architect and implement algorithms for data indexing and search capabilities. Build and maintain backend services for data storage, including work with key-value databases and synchronization. Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks. Conduct and analyze experiments on data to provide insights into system performance. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in software development. Experience with large web crawlers a plus Strong expertise in large stateful distributed systems and data processing. Proficiency in Kubernetes, and Infrastructure-as-Code concepts. Willingness and enthusiasm for trying new approaches and technologies. Ability to handle multiple tasks and adapt to changing priorities. Strong communication skills, both written and verbal. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an

AWSKubernetesRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Data Platform at OpenAI owns the foundational data stack powering critical product, research, and analytics workflows. We operate some of the largest Spark compute fleets in production; design, and build data lakes and metadata systems on Iceberg and Delta with a vision toward exabyte-scale architecture; run high throughput streaming platforms on Kafka and Flink; provide orchestration with Airflow; and support ML feature engineering tooling such as Chronon. Our mission is to deliver reliable, secure, and efficient data access at scale and accelerate intelligent, AI assisted data workflows. Join us to build and operate these core platforms that underpin OpenAI products, research, and analytics. We’re not just scaling infrastructure – we’re redefining how people interact with data. Our vision includes intelligent interfaces and AI-assisted workflows that make working with data faster, more reliable, and more intuitive. About the Role This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You’ll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You’ve supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You’re well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per wee

AWSRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About Role We are looking for an Operations Program Manager - Robotics Data Acquisition to own the day-to-day operating rhythm in our data collection facilities. You will work closely with operators, technicians, program managers, and engineers to keep rigs ready, campaigns moving, issues resolved, and performance improving. This is a hands-on operations role that requires you to be comfortable spending time on the floor, working through ambiguity, and using data to make the operation more reliable and efficient. This role is based in San Francisco, CA and requires in-person presence 5 days a week. In this role you will: Coordinate daily operations readiness across workstations, operators, materials. Track core operating metrics including utilization, cycle time, throughput, downtime, operator productivity, and data quality. Identify bottlenecks through workflow analysis, time studies, and capacity modeling, then drive practical fixes. Execute the rollout of new hardware, sensors, tools, and process changes with Engineering, Operations, Facilities, Supply Chain, and Safety. Identify equipment readiness issues and coordinate with technical support to keep workstations, and test equipment calibrated, configured, maintained, and ready for rollouts and evaluations. Lead root cause analysis for recurring operational issues and follow through on corrective actions. Provide operation input to create and maintain SOPs, work instructions, training materials, and process controls. Identify and flag resource constraints and manage issue escala

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

Overview: The Data Acquisition team within the Foundations organization at OpenAI is responsible for all aspects of data collection to support our model training operations. Our team manages web crawling and GPTBot services and works closely with Data Processing, Architecture, and Scaling teams. We are looking for a skilled Full-Stack Engineer to join our Data Acquisition team to build and optimize the interfaces and tools that power our data infrastructure. Responsibilities: Develop and maintain full-stack applications that support data acquisition, including internal tools and dashboards. Collaborate closely with cross-functional teams, including Data Processing, Architecture, and Scaling, to ensure seamless data ingestion and workflow management. Design and implement APIs to facilitate data interactions between internal services and external data sources. Enhance user experience by developing intuitive web-based interfaces for managing and monitoring data pipelines. Optimize backend services for performance, scalability, and security in a distributed computing environment. Work with legal and compliance teams to ensure our data acquisition processes adhere to privacy regulations and best practices. Deploy and maintain infrastructure using Kubernetes and Infrastructure-as-Code (IaC) methodologies. Analyze system performance, conduct experiments, and improve data workflows to maximize efficiency. Qualifications: BS/MS/PhD in Computer Science or a related field. 4+ years of industry experience in full-stack development. Proficiency in frontend frameworks (React, Vue, or similar) and backend technologies such as Python, Node.js, or Go. Strong expertise in RESTful APIs, GraphQL, and database design (SQL and NoSQL). Experience building data-intensive applications that handle large-scale datasets. Familiarity with cloud platforms (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker). Prior experience with web crawling and large-scale data processing is a

PythonReactNode.jsVue
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About The Team The Data Understanding team is responsible for creating the high quality datasets and their quantized representation for OpenAI. This includes synthesizing data, building VQ representations, and processing, filtering, deduplication, quality control, and tokenization so it can be used effectively in big model training runs. About The Role We're looking to advance how OpenAI builds and understands pretraining data at scale. You'll treat data quality and curation as core research problems: developing new methods to select, combine, and transform data; creating datasets that improve model capabilities; and designing rigorous experiments to understand how data choices and interventions affect model learning and downstream behavior. You'll work closely with frontier models and web-scale data to build evidence for which approaches work and why, then translate successful research into scalable data processing pipelines We Expect You To Have a strong track record of new or improved ML ideas, through publications, projects, or applied research. Own and drive a research agenda, from choosing the right problems to carrying long-running work through to impact. Be excited by OpenAI’s empirical, collaborative approach to research. Nice To Have Thoughtfulness about AI’s impact, including privacy, provenance, and data quality. Experience building high-performance deep learning or large-scale data processing systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The IT and Security organization builds the systems, data foundations, and automation that help OpenAI operate securely and reliably at scale. We support critical domains across identity, access, infrastructure security, enterprise systems, and internal productivity. As OpenAI grows, audit readiness and control assurance increasingly depend on reliable data: accurate system inventories, access populations, change records, configuration state, exception signals, and evidence generated directly from source systems. Our goal is to move beyond manual evidence collection and build scalable data products, automated validation, and continuous control monitoring that make security and IT controls measurable, repeatable, and defensible. About the Role We are looking for an IT Controls Data Engineer to build the data infrastructure that powers audit readiness, IT controls, evidence automation, and continuous control monitoring. In this role, you will design and maintain the pipelines, datasets, models, validation logic, dashboards, and evidence exports that make IT controls measurable, repeatable, and defensible. You will work across Security, IT, Infrastructure, Engineering, Finance Risk Management, and auditors to turn complex system behavior into reliable control data products. This is a technical builder role. The ideal candidate is strong in data engineering and analytics engineering, comfortable working with enterprise and security system data, and able to explain data lineage, source-system behavior, and control logic clearly to technical and audit stakeholders. You’ll be responsible for Building reliable data pipelines, models, and datasets for IT controls, including access, identity, configuration, change, ticketing, exception, and evidence data. Creating data quality, lineage, reconciliation, and completeness checks that make control data defensible for SOX and other audit use cases. Designing automated evidence generation workflows that produce compl

PythonSQLAWSAzure
O
📍 Seattle, Washington, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team Online Data builds and operates Habitat, the single product surface of Online Data and the system of record for OpenAI’s online user data. As OpenAI’s scale and product requirements evolve, Habitat is becoming a full-stack, one-size-fits-most database platform with end-to-end ownership of: Provisioning and developer experience APIs and guardrails Scaling, performance, and reliability Data movement, caching, routing, and placement Privacy enforcement and access control Change Data Capture (CDC) as a first-class primitive The foundation for future storage backends You’ll work on the core online database platform behind OpenAI’s products, building and operating Habitat services that handle high-QPS, latency-sensitive workloads across regions. You’ll partner closely with internal platform and product teams to ship safe, reliable systems, then push them to be faster and more cost-efficient through better caching, routing, observability, and operational tooling. This is a critical role for engineers who like owning hard distributed-systems problems end to end and sweating the details from p99 latency to production operations at massive scale. In this role, you will Design and build core abstractions spanning storage, caching, routing, CDC, and privacy enforcement Own a major surface area end to end, from product and API design to operational excellence Improve latency, correctness, and cost efficiency for real production workloads at massive scale Build strong instrumentation, debugging workflows, and developer-first tooling Collaborate closely with internal product and infrastructure teams to understand requirements and ship pragmatic solutions Participate in an on-call rotation and raise the bar on reliability while aggressively improving performance and usability You might thrive in this role if you have A strong track record building and operating high-scale backend or data-intensive distributed systems in production Excellent systems judgment and the a

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About The Team The Data Understanding team is responsible for creating the high quality datasets and their quantized representation for OpenAI. This includes synthesizing multimodal data, building VQ representations, and processing, filtering, deduplication, quality control, and tokenization so it can be used effectively in big model training runs. About The Role We’re looking to advance how OpenAI prepares, curates, synthesizes and understands multimodal data at scale. You’ll work on research and production problems like synthesizing multimodal content (images, audio, and video) and their supervisions, improving noisy data pipelines, building better quality filters, using models to automate data prep, and measuring whether changes in the dataset improve model performance. We Expect You To Have a strong track record of new or improved ML ideas, through publications, projects, or applied research. Own and drive a research agenda, from choosing the right multimodal data problems to carrying long-running work through to impact. Be excited by OpenAI’s empirical, collaborative approach to research. Nice To Have Experience with multimodal learning, audio, vision, video, synthetic data, or data-centric ML. Thoughtfulness about AI’s impact, including privacy, provenance, and data quality. Experience building high-performance deep learning or large-scale data processing systems. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of

AWSRestAIRust
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Data Clean Rooms team is Leading the market shift from traditional 2-party data sharing to multi-party collaboration hubs . Our vision is to provide a seamless, "safe-room" environment where enterprises can collaborate on shared datasets while maintaining absolute governance. We ensure that no party can exfiltrate another's underlying content, even while running complex joint workloads and getting high-value results. You will join a fast-paced, collaborative team of engineers on a journey to provide customers with an integrated set of innovative, AI-enabled capabilities to analyze data in a privacy-preserving way. You will have a real opportunity to impact and shape the future of secure data collaboration at Snowflake. AS A SOFTWARE ENGINEER IN DATA CLEAN ROOMS, YOU WILL: Architect and build highly scalable infrastructure that enables secure, multi-party collaboration. Design and implement core clean room features and services, intelligent agents, and robust developer APIs to expand platform capabilities and support custom AI/ML workflows. Partner closely with Product Management and cross-functional teams to drive complex projects from ideation and system design through to production deployment. Mentor peers and foster a warm, supportive culture of innovation, cross-tea

PythonJavaAIGo
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake Horizon Catalog is the context and governance layer of the Snowflake AI Data Cloud. AI is fundamentally transforming how enterprises monitor and trust their data: instead of manually configuring thresholds and triaging failures, AI-powered anomaly detection fires on freshness and volume deviations automatically, agentic root cause analysis traces quality failures back to their upstream source in minutes, and intelligent lineage surfaces the blast radius of any schema change before it reaches a dashboard. Snowflake's native Data Observability capabilities — Data Metric Functions, graphical column-level lineage, anomaly detection, and AI-powered root cause analysis — are the foundation of this next-generation trust layer. We are seeking a Senior Product Manager to own the Data Observability product area, spanning Data Quality, End-to-end Lineage, and Root Cause Analysis. You will define the strategy for how Snowflake competes and wins against best-in-class observability platforms, drive the AI-first transformation of how enterprises detect and resolve data quality failures, and lead a high-performing cross-functional team to deliver capabilities that make enterprise data pipelines self-healing and self-explaining at scale. AS A SENIOR PRODUCT MANAGER – DATA OBSERVAB

ReactAIGoRust
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Observe by Snowflake is an AI-powered observability platform built on the Snowflake AI Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lakehouse using open formats like Apache Iceberg — at dramatically lower cost. A dynamic Context Graph and chat-based AI SRE provide rich context and automated workflows so teams can move from detection to root cause and resolution 10x faster. Leading engineering teams at companies like Capital One, Topgolf, and Dialpad rely on Observe to troubleshoot hundreds of terabytes of telemetry daily while maintaining reliability at enterprise scale. As part of Snowflake, Observe combines startup-style ownership and velocity with the global reach, operational excellence, and ecosystem of one of the world's leading data platforms. We are hiring a Senior Software Engineer for the Observe Data Management team. This team owns the core pipelines that ingest and process over 1 petabyte of telemetry data per day — the foundational infrastructure powering Observe's entire observability stack. You'll be working at the intersection of massive scale, open-source innovation, and real-world reliability challenges for enterprise customers around the globe. AS A SENIOR SOFTWARE ENGINEER - OBSERVE

AWSAzureAIC++
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. Responsible for understanding the processes used to assess and mitigate risks related to the introduction of new package technologies so that you can improve efficiency and effectiveness by implementing automation, data analysis and machine learning/AI solutions. Set up data bases, develop data ingestion pipelines, and implement automated data analysis and reporting tools. Support problem solving and risk assessment for internal or customer quality issue by pulling and analyzing data. Contribute to the advancement of technology at Micron through mentoring, publishing technical papers (internal and external), and developing innovative solutions to challenging problems. Implement Automation, Data Analysis, and AI Solutions. Collaborate with Engineering teams to Map Package DDQA processes and data streams. Set up and optimize databases and develop solutions to improve efficiency and effectiveness. Understand the needs of internal customers and develop solutions. Support Problem Solving and Risk Assessment for Quality Issues. Pull relevant product information, manufacturing data, and reliability data based on given problem statements. Determine the appropriate dataset and treatment required to answer questions posed by problem solving teams. This could include producing data visualizations, machine learning models, statistical inferences, and web applications. Provide recommendations about root cause findings and product risk, based on data analysis. Collaboratively Communicate Findings and Best Practices. Share best practices with global teams to enable a cultur

JavaScriptPythonJavaReact
C
📍 Tampa Florida United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

At Citi Services - Global Trade and Working Capital Solutions (TWCS) Technology Organization, we are on a mission to harness the power of data to drive innovation, create exceptional customer experiences, and solve complex business challenges. Our data team is at the heart of this mission, building the scalable and resilient infrastructure that turns data into our most asset. We are a passionate, collaborative group dedicated to pushing the boundaries of what's possible. The Opportunity We are seeking an Engineering Director to join our development team. The ideal candidate is a seasoned technologist with extensive experience in building and delivering data & AI solutions for business functions. This individual will be directly and fully accountable for solution delivery and should have a history of creating strategic technology architecture roadmaps aligned with business outcomes. The successful candidate will define and execute the technology roadmap for the Data & AI portfolio of TWCS, providing strategic direction and critical input into technology decisions to ensure a scalable buildout. This role involves building strong relationships with senior business and technology partners, driving agile execution, and leading a team of expert engineers to deliver with velocity and quality. Applicants should demonstrate exceptional technical acumen, a strong data engineering background, and a proven ability to lead and provide technical direction to high-performing engineering teams. A proven expertise in Data Engineering, Data Analytics, and AI Engineering delivery at scale is essential. Responsibilities: Manage/develop multiple teams of professionals to accomplish established goals and conduct personnel duties for team (e.g. performance evaluations, hiring and disciplinary actions) as well as ensure team adheres to best practices and processes Develop vision for team aroun

PythonJavaAWSAzure
P
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -72.3%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. Plaid’s Partnerships team develops strategic relationships with technology platforms and other ecosystem partners that expand the reach and impact of Plaid’s products. This role will focus on data and product partnerships that power Plaid’s Credit products and insights. Plaid’s Credit team is building the future of lending by making cash flow data as widely used and trusted as traditional credit data. Our mission is to expand access to more affordable credit by giving lenders real-time financial insights that improve risk assessment and decision-making at scale. In this role, you will build and manage the external data partnerships that power Plaid’s Credit products. The partner ecosystem is broad and includes traditional credit data providers, alternative data providers, financial infrastructure companies, and other sources of differentiated data and insights. You will work closely with Product and Engineering to understand where new data can improve existing products or unlock entirely new capabilities. You will identify and prioritize partners, develop creative commercial structures, and negotiate complex agreements that balance product, economic, and strategic considerations.

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.

AWSAzureRestAI
🔔

Get new data entry specialist jobs in United States by email

Daily job updates · Unsubscribe anytime