AI Research Engineer/Scientist — US, California, Santa Clara. Apply via Workday.
Jobiba hiring network
Ai Research Scientist Jobs
10,000 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current ai research scientist jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Research Scientist – AI in Computed Tomography Image Formation (all genders) — Hamburg. Apply via Workday.
About the Team The Recursive Self-Improvement (RSI) team works across research, engineering, product, and infrastructure to build AI systems that accelerate and ultimately conduct high-quality research at OpenAI. We work to automate real research workflows and improve research productivity by building systems and feedback loops, designing evaluations, and training models to develop missing capabilities. Our work spans the full lifecycle of model training, evaluation, and deployment to help researchers move faster and tackle increasingly ambitious problems. About the Role We’re hiring research scientists , research engineers , and AI systems engineers to work on automating research at OpenAI. This role is based in San Francisco, CA. In this role, you will: Design evaluations for research judgment, hypothesis generation and testing, and long-horizon experiment execution. Turn real research workflows and model failures into data and evaluation flywheels. Improve model research capabilities through agent harnesses, synthetic data, RL environments, and model training. Build and maintain safe, reliable integrations between our models and OpenAI’s research infrastructure. Develop research agents, experiment-orchestration systems, and sandboxed runtimes that support real research workflows. Create metrics and economic models to understand RSI’s current and future effects on research productivity, model capabilities, and the safety of internal deployments. This is a high-ownership role for researchers and engineers who thrive in ambiguity, move fluidly between research and implementation, and turn emerging opportunities into rigorous, reliable, scalable results. You might thrive in this role if you: Have research or engineering experience across LLM training, model evaluations, agent systems, synthetic data, research infrastructure, or large-scale distributed systems. Are a strong generalist who can move between open-ended research and practical implementation, turning ambig
About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! We’re looking for a Research Scientist to advance how high-quality data and environments for AI agents are created. You’ll build and optimize pipelines that combine real-world data, automated generation, and human expert input. Working with domain experts, academic partners, customers, and our product and engineering teams, you’ll scale these pipelines to target frontier model performance gaps and expand data and environment diversity. Your work will amplify human knowledge and judgement, enabling experts to create and refine data and agentic environments that strengthens Snorkel’s position as the frontier data lab. This role is ideal for someone who wants to advance frontier AI through data and environment creation and enjoys turning research into reusable, scalable systems. Location: San Francisco, New York, OR REMOTE Main Responsibilities Design, implement, and optimize reusable pipelines that combine AI capabilities with expert judgment to accelerate data and agentic environment creation. Design and run rigorous experiments to validate proof-of-concept approaches, measure their impact on data quality, pipeline efficiency, and model performance, and communic
Scale works with the industry’s leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling). This role will focus on optimizing data curation and eval to enhance LLM capabilities in both text and multimodal modalities. In this role, you will develop novel methods to improve the alignment and generalization of large-scale generative models. You will collaborate with researchers and engineers to define best practices in data-driven AI development. You will also partner with top foundation model labs to provide both technical and strategic input on the development of the next generation of generative AI models. You will: Research and develop novel post-training techniques, including SFT, RLHF, and reward modeling, to enhance LLM core capabilities in both text and multimodal modalities. Design and experiment new approaches to preference optimization. Analyze model behavior, identify weaknesses, and propose solutions for bias mitigation and model robustness. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning. Excellent written and verbal communication skills Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals Previous experience in a customer facing role. Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined du
Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities. In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models. You will: Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA. Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them. Publish research findings in top-tier AI conferences. Ideally you’d have: Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field. Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning. Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development. Excellent written and verbal communication skills. Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals. Previous experience in a customer facing r
ROLE SUMMARY We are seeking a highly motivated individual to join our high content imaging lab within the Discovery Biology and Pharmacology (DBP) group at the vibrant Groton campus in Connecticut. DBP is responsible for hit-identification, lead optimization, and molecular characterization of our small molecules and has teams focused on high-throughput screening, DNA-encoded libraries, pharmacology, protein homeostasis platforms, cellular models, functional genomics, and high-content imaging. The group is highly integrated with other groups in Medicine Design including Chemists, Structural Biologists, and Computational Scientists, and supports the small molecule portfolio across multiple therapeutic areas. The individual in this role will bring in rich experience in high content imaging and high throughput flow cytometry, and apply them to support our diverse small molecule portfolio spanning various therapeutic areas. ROLE RESPONSIBILITIES High-content imaging assay design, development, and optimization: Apply a broad range of imaging and flow cytometry assay technologies to address project needs. Independently develop, optimize, and troubleshoot high-content imaging and flow cytometry assays. Imaging data analysis and interpretation: Design and implement advanced image-analysis workflows, build complex analysis algorithms, and streamline data processing to support medium- to high-throughput screening. Translate imaging data into clear, actionable biological insights. Imaging infrastructure management and continuous improvement: Partner with team experts to support high-content imaging and flow cytometry instrumentation, implement software solutions that enable advanced imaging applications and data analysis, and continuously improve ima
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must continue to develop new evaluation techniques that accurately reflect what models are already capable of, as well as set the agenda for what future models should be capable of. In this role, you are responsible for creating these next-generation evaluation methods and infrastructure to measure LLM progress. As a Senior Research Scientist, Model Evaluation, you will: Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish. Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations. Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. Build scalable and reusable tools for digging into model performance. You may be a good fit if: You enjoy rapidly building prototypes that demonstrate the boundaries of what LLMs are capable of, and you have developed res
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake's Data Engineering organization builds the platform that ingests, transforms, and stores data for modern lakehouse architectures — powering billions of queries, DML, and DDL operations with industry-leading price-performance. We lead the industry's shift to open data lakes through our work on Iceberg and Polaris, and we deliver capabilities like Snowpark, Dynamic Tables, cross-region replication, time travel, and zero-copy cloning at enterprise scale. We are investing in a new line of applied research — building toward verified data infrastructure and trustworthy data systems — that brings formal methods, automated reasoning, and modern AI techniques to bear on the hardest problems in our distributed systems and developer tooling. The goal is to improve correctness, reliability, and engineering velocity at a scale very few platforms operate at. We're hiring at both the Staff and Principal level; we'll calibrate the offer to the candidate's experience and scope of impact. What you'll do Lead research projects that apply formal methods, program analysis, automated reasoning, and AI-driven techniques (including code generation and modeling) to real problems in our cloud data platform. Translate research ideas into prototypes, then into shipped capabilities that move
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Work Develop and scale state-of-the-art generative models—especially diffusion architectures, flow-matching techniques, and energy-based models —for autonomous plan generation. Build generative models with foundation models. Leverage large language models and world foundation models for reasoning, decision making and multi-modality generation. Optimize generative models using reinforcement learning to improve interactive reasoning. Explore reward modeling/learned verifier using generative models. Explore joint prediction and planning and self-play. Leverage generative models for active learning and world modeling. Develop controllable generative models to guide the generation process towards desired goals, conditions and rewards. Collaborate across autonomy teams while developing holistic solutions to top autonomy challenges. Understand issues, propose ideas, prioritize work and develop solutions to solve them, evaluate
About the Team OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make better, evidence-based talent decisions. About the Role As a People Research Scientist, you will bring deep expertise in research design, measurement, experimentation, and applied data science to OpenAI’s most important People programs. You will design studies, evaluate people processes, and help leaders better empower employees, strengthen organizational systems, and deliver exceptional employee experiences. This is a high-ownership individual contributor role combining hands-on research, methodological leadership, and scalable people science capabilities. We’re looking for an experienced researcher who can turn ambiguous People questions into rigorous designs, validated insights, and actionable recommendations. This role is based in San Francisco, CA or Mountain View, CA, with occasional travel to our San Francisco office. What You’ll Do: Design rigorous research and evaluation strategies for recruiting, organizational health, manager effectiveness, employee experience, and talent outcomes. Apply advanced statistical modeling, machine learning, and research methods to inform program design, evaluate effectiveness, and quantify business impact. Partner with People Operations, data engineering, and people systems teams to define data requirements, improve data quality, establish documentation standards, and ensure research datasets are governed, reproducible, and privacy-preserving. Build scalable people science infrastructure, including self-service agentic tools, automated validation workflows, reusable research datasets and analytical pipelines. Develop research playbooks that establish rigorous standards for study design, measurement, validation, and documentation, enabling high-quality, repeatable, and scalable research across the organization. Communicate findings through c
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role The mandate of the prediction team is to use advanced machine learning techniques to improve the behavior of the Nuro Driver. As a key member of the Prediction and Smart Agents team, you will focus on building state-of-the-art models for predicting the behavior of surrounding traffic. These models are crucial for our autonomous system, as they will be deployed onboard as part of our planning stack and used offboard for realistic closed-loop simulation. You will explore novel machine learning methods to solve challenging real-world problems in autonomous driving. This work includes using generative sequence modeling approaches for robustly predicting complex, interactive traffic situations. It requires deep reasoning about the intentions of other road users and how their behaviors influence safe and correct driving decisions. You will also use different input modalities, including End-to-End (E2E) approaches, for predicting
About the Team The Health team, within OpenAI’s broader Personal AGI organization, has a mission to ensure AGI improves health for all humanity. Improving human health will be one of the defining impacts of AGI. Hundreds of millions of people already turn to ChatGPT for questions about their health and millions of clinicians use it weekly to support care delivery. Increasingly capable models create an opportunity to make high-quality medical intelligence more accessible across patients and clinicians—raising the floor of human health—and accelerate the new capabilities and scientific advances that raise the ceiling of human health. Our job is to make those benefits real. We work across the full model stack—pretraining, midtraining, reinforcement learning, post-training, evaluations, harnessing, and deployment—and connect that research to the patients, clinicians, and real-world outcomes we aim to improve. About the Role We’re looking for an exceptional, hands-on researcher who wants to build frontier health capabilities and turn them into impact at scale. This is a role for someone who can take an important, underdefined problem from 0→1: identify the right bet, build what’s needed to test it, and drive it all the way to a measurable improvement in the models and products we actually ship. We’re especially excited about two kinds of people: researchers with the technical depth to move the frontier in pretraining, reinforcement learning (RL) / post-training, or evals; and researchers with real depth in developing frontier biomedical AI capabilities. Prior experience in healthcare is helpful but not required. Research excellence, velocity, ownership, and alignment with the mission are most important to us. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own a high-leverage research direction end to end—from deciding which problem matters and h
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the Role In this role, you will be a member of the Perception & Behavior team, leveraging the cutting edge of machine learning research to solve challenging real-world robotics problems. This role is focused on bringing advancements in the field of ML and large-scale learning to the AV domain, moving towards a more end-to-end autonomous driving system. This role requires working with and developing large models for perception and behavior, keeping up-to-date and experimenting with state-of-the-art architectures quickly and efficiently, collaborating with other teams to determine data and infrastructure support needs, and working to improve model optimization and inference speeds. You will use your applied research skills to think through the creation and deployment of these ML models on our autonomous vehicles while working alongside talented researchers in the field. If you love solving fundamental AI challenges and deploying your s
Who We Are Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides. Founded in 2016, Nuro is a physical AI company developing Level 4 autonomous driving technology for a wide range of vehicles, use cases, and markets. Powered by the Nuro Driver™, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles. With years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Nuro has raised over $2B in capital from Uber, NVIDIA, Google, Softbank, Fidelity, T. Rowe Price, and other leading investors. About the role In this role, you will collaborate closely with researchers and engineers on the Learned Behavior teams to tackle plan generation challenges in autonomous driving. You’ll apply state-of-the-art generative modeling techniques—ranging from cutting-edge diffusion, flow matching, energy-based models, and SoTA algorithms—in order to develop novel solutions that generate safe, comfortable, and efficient driving behaviors in the most challenging real world situations. Beyond core research, you’ll own the end-to-end lifecycle of your models, productizing them for robust, real-world autonomous driving deployments on a global scale. About the Work Develop and scale state-of-the-art generative models—especially diffusion architectures, flow-matching techniques, and energy-based models —for autonomous plan generation. Build generative models with foundation models. Leverage large language models and world foundation models for r
Get new ai research scientist jobs by email
Daily job updates · Unsubscribe anytime