At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. At Snowflake, we are building a high-impact team to help the world's most innovative companies unlock the power of AI. As a Senior Forward Deployed Engineer, Applied AI on our Cortex AI team, you will be a hands-on technical leader and trusted partner to our most strategic customers. You will own the end-to-end delivery of enterprise AI programs, leading a team of 2–4 engineers while staying deeply technical yourself. You will set the technical direction for your customer engagements, mentor your team, and serve as the senior technical voice at the intersection of product, engineering, and customer success. IN THIS ROLE AT SNOWFLAKE, YOU WILL: Lead Customer Programs : Own the full lifecycle of complex, multi-engineer AI engagements – from scoping and architecture through deployment, monitoring, and handoff. Be accountable for delivery quality and customer outcomes for the projects you lead. Own AI Quality : Define what "good" means for each engagement. Translate ambiguous customer goals into measurable quality metrics, evaluation frameworks, and golden datasets – then run systematic eval loops to hill-climb on agent quality, catch regressions before customers do, and continuously raise the bar on accuracy, faithfulness, and safety. Set the standard for how the team measures
Jobiba hiring network
Test And Evaluation Lab Tech Jobs
1,734 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current test and evaluation lab tech jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team The Proactivity Research team, within OpenAI’s broader Personal AGI team, is focused on making our models in ChatGPT and future potential products proactive in ways that are truly useful. We're laying the technical foundations for AI that can anticipate what users need in real time, adapt as their goals and preferences shift, and build a deeper, evolving understanding of the person it's helping. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models’ personalization and agentic capabilities. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a highly personalized, collaborative, and proactive assistant. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda to improve the proactivity and ability of our models to further user goals. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the other research and product teams to influence the shape of technical solutions in the product You might thrive in this role if you: Have a deep understanding of machine learning and machine learning applications. Have a working knowledge of LLM post-training and evaluation approaches Are passionate about, or have experience thinking about, personalization and enabling users to achieve their goals Are comfortable diving into a large ML codebase to debug. Thrive in a dynamic and technically complex environment. About OpenAI
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence. In this role you will Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows Partner closely with applied AI engineers and product leaders to define what “good” looks like and translate it into measurable criteria Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring You’ll love this job if you Care deeply about correctness, rigor, and measurement in AI systems Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team Thrive in cross-functional environments and enjoy influencing model design through better evaluation Qualifications Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learni
Scale Labs, Research Scientist — Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Agent Robustness you will work on the fundamental challenges of building AI agents that are safe and aligned with humans. For example, you might: Research the science of AI agent capabilities with a focus on how they relate to safety, risk factors, and methodologies for benchmarking them; Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured to do so by users or tricked into doing so by elements of their environment; Design and build exploits and mitigations for new and unique failure modes that arise as AI agents gain affordances like coding, web browsing, and computer use; Characterize and design mitigations for potential failure modes or broader risks of systems involving multiple interacting AI agents. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and leveraging agent scaffolding, designing evaluation harnesses, an
MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn healthtech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really is : We are seeking a detail-oriented and technically strong AI QA Engineer to ensure the quality, reliability, and performance of Large Language Model (LLM)-based systems. In this role, you will be responsible for designing and executing test strategies, validating model outputs, and building evaluation frameworks to enhance the accuracy, safety, and overall performance of AI-driven applications.We would particularly value candidates who have hands-on experience in developing evaluation frameworks (evals) for AI systems, along with strong expertise in comprehensive system testing and quality assurance practices.You are responsible for making MeltPlan work in the real world. What You'll Do: Design, develop, and execute evaluation frameworks (Evals) for Large Language Models (LLMs) and AI syst
About the Team The Human Data team at OpenAI is responsible for identifying and mitigating risks in advanced AI systems by designing evaluations, surfacing vulnerabilities, and collaborating closely with researchers to strengthen model reliability and public trust. About the Role As a Research Program Manager, you will lead initiatives that test the safety and robustness of OpenAI’s models through creative experimentation and structured evaluation. You’ll coordinate efforts across research and engineering teams to transform ambiguous risks into concrete research programs and influence future model development and deployment. We’re looking for people who are technically savvy, comfortable with ambiguity, and excited about shaping the future of safe AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead programs that explore unexpected model behaviors and identify failure modes. Translate vague or emergent risk signals into clear priorities and actionable research plans. Design and run creative evaluations, experiments, and red-teaming campaigns. Collaborate with research, product, and deployment teams to integrate findings into model training and deployment cycles. Develop repeatable systems for tracking model performance and understanding emerging behavior patterns. You might thrive in this role if you: Have strong experience in technical program management, with excellent organizational and communication skills. Are familiar with large language models, prompt engineering, or model evaluation techniques. Are comfortable managing fast-paced, high-uncertainty projects and shaping them from the ground up. Are creative and resourceful in devising new methods for testing model behavior and performance. Can effectively coordinate across technical and non-technical stakeholders to drive alignment and execution. About OpenAI OpenAI is an AI resear
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer for our Frontier Security AI team. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers — products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. In this role, you will lead the design and development of our Agentic Harness and agent evaluation platform, working across product, infrastructure, applied AI, security, and modeling teams to take new capabilities from prototype to dependable customer value. AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. Convert ambiguous reports such as "the agent feels worse" into measurable failure modes, reproducible tests, and durable fixes. Analyze production agent trajectories to identif
About the Role In this role, you’ll lead and evolve the design system foundations for ChatGPT. You’ll make nuanced design judgment explicit and usable across product teams, establishing the principles, patterns, and standards that enable coherent, high-quality experiences at scale. A paramount part of this role is raising OpenAI’s craft bar. This work extends beyond maintaining a component library: you’ll define what excellent product design looks and feels like, establish clear guidance for choosing and applying interface patterns, and ensure that new ChatGPT experiences strengthen a coherent system across platforms. Working closely with research, engineering, product, and design, you’ll create a continuous learning loop between design principles, product experiences, evaluation, and evolving technology—shaping a system that becomes more capable and expressive as ChatGPT grows. This role is based in our San Francisco HQ. We offer relocation assistance to new employees. In this role you will: Define an exceptionally high standard for craft and taste, translating that standard into principles, guidance, and examples that elevate the quality of ChatGPT. Own and evolve ChatGPT’s design foundations, including typography, sizing, spacing, layout, motion, accessibility, and responsive behavior. Establish a shared source of truth across platforms and improve the systems connecting tokens, components, patterns, widgets, learning experiences, and app-like product experiences. Build a comprehensive design systems framework with clear, opinionated guidance about when and how product teams should use different interface patterns. Translate design judgment into durable standards and use prototyping and evaluation workflows to test experiences, identify recurring issues, and improve the system. Identify promising patterns across emerging product experiences and turn them into reusable, documented foundations. Partner with research, engineering, product, and design to evolve the d
About the Team The Future of Computing Research team is an applied research team within the Consumer Devices group focused on developing new methods, models, and evaluation frameworks that support our vision for the future of computing. We work at the frontier of multimodal AI, helping turn emerging model capabilities into product experiences that are useful, delightful, and worthy of long-term trust. Our work explores a new class of AI systems that can learn over time, adapt to individuals, and support people in the flow of daily life. This includes long-term memory, user modeling, and personalization systems that are aligned not just with immediate satisfaction, but with a person’s broader goals, values, and well-being. We work closely across research, engineering, design, product, and safety to define what it means to build AI systems that know you over time, act at the right moment, and help in ways that are context-aware, respectful, and demonstrably beneficial. About the Role We are looking for a Research Engineer / Scientist to join the Future of Computing Research team to work on RLHF and post-training for personalized, multimodal AI systems. This role will focus on building the learning and evaluation foundations that help models become more context-aware, adaptive, and useful over time. You will work on problems such as reward modeling, preference learning, long-horizon evaluation, and policy improvement for systems that must make high-quality behavioral decisions in realistic user settings. The work is deeply product-grounded: success is not just higher benchmark performance, but better model behavior in real-world use. The ideal candidate is excited about pushing beyond one-turn assistant behavior toward systems that improve through feedback, learn from richer signals, and are trained against meaningful notions of user value. Internally, that maps closely to the need for careful reward design, feedback loops, and evaluation frameworks that test whether i
Human Data Quality Analyst, AI Business Prolific Prolific isn’t just enabling AI innovation – we’re redefining it. While foundational AI technologies are becoming commoditized, Prolific’s human data infrastructure provides the high-quality, diverse data required to train the next generation of AI models. Through our platform, we empower researchers and companies to access a global, ethically curated participant base, ensuring cutting-edge AI research and training grounded in inclusivity and precision. The Role Prolific provides the human data that powers the next generation of AI models, working with frontier labs to capture the complex human judgments researchers need to train, evaluate and improve them. As a Human Data Quality Analyst, you'll be on the front line of making sure that data captures the right signal and is genuinely good enough to do its job. This isn’t traditional, back-office QA. You’ll be doing real analytical work: digging into datasets, identifying patterns and failure modes, investigating why quality has shifted, and turning complex findings into clear insights that help us improve how data is collected, reviewed and delivered. You'll spend real time reading annotations closely, but that is how you gather evidence, not what you produce. What you produce is analysis, practical recommendations and better quality controls. You'll get hands-on exposure to human data, annotation, machine learning pipelines and AI evaluation, working alongside Quality, Engineering, Operations and Delivery on new and evolving problems. There won't always be an established playbook. You'll be guided by our quality engineers, but you'll also need to run your own analysis, test your assumptions and recognise when you need input. It is a role with a steep learning curve from day one and a strong opportunity for someone early in their career to build deep, practical experience in a fast-moving area of AI. What You’ll Be Doing Run day-to-day qualit
About the Team OpenAI’s Forward Deployed Engineering team partners with leading semiconductor companies to deploy production-grade AI systems across the entire chip design lifecycle: design, verification, and physical design. We operate at the intersection of customer delivery and core platform development, embedding deeply with customers to translate frontier model capabilities into systems that materially improve engineering workflows and accelerate innovation. Our work turns early, high-touch deployments into repeatable solution patterns, reference architectures, and evaluation practices that scale across the semiconductor ecosystem. About the Role We are seeking a highly skilled Physical Design Engineer to join our semiconductor-focused Forward Deployed Engineering team. This is a senior IC role that will begin with a strong emphasis on physical design expertise, technical judgment, advisory leverage, and customer credibility, with the expectation that the person will grow into a broader Forward Deployed Engineering role over time. In the near term, you will serve as the team’s physical design SME across semiconductor deployments: helping FDEs, Product, and Research understand backend implementation workflows, pressure-test AI-assisted solution ideas against real physical design constraints, and raise the quality of our customer-facing technical work. You will help the broader team build fluency in implementation flows, EDA tooling, signoff methodology, and the trade-offs that shape physical design decisions in practice. Over time, we expect this role to expand beyond SME support into broader FDE ownership: partnering directly with customers, shaping deployment strategy, building and iterating production-grade AI systems, driving technical workstreams, and helping turn high-touch semiconductor deployments into repeatable solutions. This is a strong fit for someone who brings deep physical design expertise today and is excited to grow into a customer-facing, syst
About the Team The Personalization-Memory team, within OpenAI's broader Personal AGI organization, is focused on developing agents that can learn from prior interactions in order to become more helpful and efficient over time. We build general-purpose memory and personalization capabilities that transfer across ChatGPT and other agentic products, and we collaborate with applied engineering on the product surfaces that allow users to interact with memory. About the Role As a Research Engineer / Research Scientist on the Personalization-Memory team, you will research and develop improvements to memory usage and personalization in OpenAI's frontier models. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a truly personalized ChatGPT. We're looking for individuals who have a background in frontier model post-training, are able to iterate quickly, and who are passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda for improving memory use and personalization in frontier models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the research and product teams to influence the shape of technical solutions in the product. You might thrive in this role if you: Are passionate about personalization and building personalized assistants. Have experience working with user signals and human data to turn feedback into reliable signals for training and evaluation. Have a deep understanding of frontier model post-training and machine learning applications. Value principled approaches and research craftsmanship. Are comfortable diving into a lar
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for a driven and collaborative IP attorney to join Snowflake's growing Intellectual Property team. Reporting to the Director, Assistant General Counsel – Intellectual Property, this role will be a core member of a lean, high-impact IP team where you will have direct ownership over meaningful work from day one. You will manage Snowflake's patent portfolio, lead patent prosecution strategy, and support pre-litigation and IP litigation matters. This is an exceptional opportunity for a patent attorney who brings both prosecution and litigation experience, thrives in a fast-moving technology environment, and is eager to grow their practice into adjacent areas such as open source. This is a hybrid role based in Menlo Park, CA or Dublin, CA with a preference for Menlo Park where the majority of the IP team is based. RESPONSIBILITIES Manage Snowflake's global patent portfolio, including overseeing prosecution strategy, outside counsel relationships, and budget Lead and continuously improve Snowflake's internal patent program, including inventor education, invention submission, and evaluation processes Review and supervise outside patent counsel on the drafting of new patent applications and prosecution strategy across Snowflake's patent portfolio, working closely wit
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . As a Staff Data Scientist on the Data Science team within the Platform group, you'll own pricing experimentation and strategy for Coinbase's Consumer & Business products. This team partners directly with Product, Engineering, and Design to turn deep analytical expertise into decisions that move the company's bottom line. You'll design and analyze pricing tests, build models that identify optimal strategies, and communicate findings to executives and cross-functional leaders. What you'll do: Own end-to-end pricing experimentation, from test design through analysis and recommendation Build and refine pricing models and evaluation frameworks that determine optimal pricing strategy across consumer products Partner with Product, Engineering, and Finance stakeholders to develop pricing vision, roadmap, and priorities Develop and maintain data pipelines and data models that power pricing analytics with production-grade craftsmanship Synthesize complex findings into clear, actionable recommendations and present them to senior leadership Required Skills and Experience: 8+ years of experience in data science with a focus on pricing experimentation, causal inference, and statistical modeling (PhD preferred, or Master's in Economics, Statistics, or related quantitative field) 5+ years directly leading pricing or experimentation workstreams, including designing A/B tests and
Senior Test & Evaluation Engineer (Test Program Req & Planning) Company: The Boeing Company Boeing Test & Evaluation (BT&E) organization is seeking a Senior Test & Evaluation Engineer (Test Program Req & Planning) to join our team in Oklahoma City, OK . This position supports a diverse portfolio of Boeing programs including the B-1, and B-52 sustainment and modernization programs. The Engineer is responsible for supporting the Test Program Manager (TPM) and other program leadership by leading the Ground Test/Flight Test (GT/FT) Integrated Product Teams (IPTs) and ensuring the integration of Boeing Test & Evaluation (BT&E) capabilities and resources to meet program commitments. This role requires close collaboration with program IPTs, Systems Engineering Integration & Test teams, Program Offices and external developmental test organizations Position Responsibilities: Develop, prepare, and coordinate comprehensive Program Integrated Test Plans (ITP), detailed Test & Evaluation (T&E) plans, and associated documentation including requirements traceability, verification metrics, and program requirements. Lead the development and standardization of test methods, procedures, and processes to ensure repeatability, productivity, and continuous improvement through lessons learned. Integrate BT&E capabilities and resources with other IPTs, Systems Engineering, government stakeholders, industry partners, suppliers, and customers to ensure seamless test execution and alignment with program objectives. Conduct testability assessments of engineering requirements and oversee disciplined cost, schedule, and configuration management to meet program baselines and commitment
Get new test and evaluation lab tech jobs by email
Daily job updates · Unsubscribe anytime