Scale Labs, Research Scientist — Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist working on Agent Robustness you will work on the fundamental challenges of building AI agents that are safe and aligned with humans. For example, you might: Research the science of AI agent capabilities with a focus on how they relate to safety, risk factors, and methodologies for benchmarking them; Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured to do so by users or tricked into doing so by elements of their environment; Design and build exploits and mitigations for new and unique failure modes that arise as AI agents gain affordances like coding, web browsing, and computer use; Characterize and design mitigations for potential failure modes or broader risks of systems involving multiple interacting AI agents. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and leveraging agent scaffolding, designing evaluation harnesses, an
Jobiba hiring network
Test Job 02 09 Jobs
1,747 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current test job 02 09 jobs. Use filters to narrow by work mode, employment type, experience and date posted.
This position is based in Vancouver, BC , within Diligent’s Technical Center of Excellence. We are currently hiring candidates who are based in or able to work from Vancouver . Software Engineer — Platform AI Service Levels: Software Engineer II Senior Software Engineer Staff Software Engineer Location: Vancouver Position Overview As a Software Engineer on Diligent's Platform AI team, you'll help design, build, and operate the core services that power AI-driven capabilities across Diligent's global product suite. You'll build secure, scalable, serverless services on AWS that translate AI research and models into commercial-quality, production-ready solutions — enabling customers to derive insights from their governance data. You'll work closely with AI researchers, product managers, and other engineering teams, owning your services end-to-end: architecture, implementation, deployment, and monitoring. The team operates with a strong AI-augmented engineering culture — using AI tools to accelerate coding, testing, debugging, and delivery — while applying sound judgment about when and how to apply them. Key Responsibilities Design and implement secure, scalable, fault-tolerant, high-performing solutions using AWS serverless technology — event-driven, highly observable, and built with infrastructure as code. Collaborate with AI researchers/engineers to translate AI and LLM capabilities into robust, production-grade services, and help other teams integrate them. Build and maintain the pipelines needed to deploy, monitor, and manage AI services at scale — observable, resilient, and cost-effective. Use AI-powered development tools (code assistants, test generation, architecture exploration) responsibly to accelerate delivery and improve quality, always validating outputs. Participate in architecture discussions and design reviews, and contribute to product design by understanding customer problems — especially where AI can offer a breakthrough solution. Work in
Here’s a summary of the role: Build cloud software that matters, grow your technical depth, and use modern AI tooling to do your best work. This is a hands-on engineering role for someone who enjoys solving product problems, writing clean code, and helping services run reliably at scale. You’ll work on secure, scalable microservices and APIs using TypeScript, AWS , and modern engineering practices. You’ll be part of a collaborative product engineering team where you can own features, contribute to design discussions, support production systems, and keep growing across backend, cloud, and AI-assisted development workflows. Here’s a breakdown of what you’ll do, not all of it, just the important stuff: Design, build, test, and improve backend services and APIs using Node.js, TypeScript, and AWS . Take ownership of well-defined features from planning through release, including code quality, deployment, and production support . Work closely with product managers, designers, and other engineers to turn requirements into practical, reliable solutions. Contribute to technical design conversations, code reviews, and engineering standards that keep the team moving well. Use AI tools to speed up research, coding, debugging, testing, and documentation, while checking outputs carefully and applying sound judgment. Help keep systems secure, observable, and maintainable by improving monitoring, reliability, and day-to-day development practices. These are the essentials you’ll need to get an interview: 3 to 5 years of professional software engineering experience building production applications in an agile environment. Strong backend development skills with Node.js and TypeScript, including experience building APIs or microservices. Experience with React or Angular in a product engineering environment. Hands-on experience with
Role Overview You’re a hands-on backend engineer who enjoys owning features end to end and working on real products that customers rely on every day. In this Software Engineer II role, you’ll help build and evolve a Third Party Risk Management SaaS platform using Laravel and PHP, designing scalable APIs and services that keep performance and reliability front and center. You’ll work in a product-focused team that owns its services from architecture and implementation through deployment, monitoring, and continuous improvement. You’ll mentor junior engineers, influence technical decisions, and use modern AI-powered tools thoughtfully to ship better code faster. If you’re looking for a mid-level role with real ownership, modern tooling, and the chance to grow your impact, this is for you. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Design, build, and maintain backend features and RESTful APIs in Laravel within a modern TALL stack environment. Own well-defined stories from implementation through deployment, monitoring, and iteration, ensuring performance and reliability. Contribute to architectural discussions and technical decisions that shape the Third Party Risk Management platform. Review code, improve test coverage, and strengthen CI/CD and engineering standards across the team. Mentor Software Engineer I colleagues through code reviews, pairing, and knowledge sharing. Use AI tools (e.g. GitHub Copilot, ChatGPT) to accelerate coding, debugging, testing, and documentation—while critically validating outputs and ensuring safe, responsible use. These are the essentials you’ll need to get an interview 3–5 years of professional software engineering experience in an agile, fast-paced environment. Strong experience with PHP and Laravel, ideally within the TALL stack (Tailwind, Alpine.js, Laravel, Livewire). Solid understanding of relational databases (MySQL or MariaDB), including data modelling and query optimisation. Experience designin
Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p
Software is eating the world, but AI is eating software. We live in unprecedented times – AI has the potential to exponentially augment human intelligence. Every person will have a personal tutor, coach, assistant, personal shopper, travel guide, and therapist throughout life. As the world adjusts to this new reality, leading platform companies are scrambling to build LLMs at billion scale, while large enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human eval and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT to get such a large headstart among competition. At Scale, our products include the Generative AI Data Engine, SGP, Donovan, and others that power the most advanced LLMs and generative models in the world through world-class RLHF, human data generation, model evaluation, safety, and alignment. The data we are producing is some of the most important work for how humanity will interact with AI. At the foundation of these products is the Platform Engineering team. In this role, you will support the design and development of shared platforms used across Scale. This includes designing our foundational data platforms and lifecycle, architecting Scale’s core cloud infrastructure and orchestration stack, and redefining how engineers develop, build, test, and deploy software at Scale. You’ll also get widespread exposure to the forefront of the AI race as Scale sees it in enterprises, startups, governments, and large tech companies. You will: Drive the design, and implementation of our foundational platforms and systems, working closely with stakeholders and internal customers to understand and refine requirements. Collaborating with cross-functional teams to define, design, and deliver new features. Proactively identifying opportunities for, and driving improvements to, current p
Scale Labs, Research Scientist — Frontier Risk Evaluations As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the gap between AI research and global policymakers to make informed, scientific decisions about AI risks and capabilities. Our research tackles the hardest problems in agent robustness, AI control protocols, and AI risk evaluations to help governments, industry, and the public understand and mitigate AI risk while maximizing AI adoption. This team collaborates broadly across industry, the public sector, and academia and regularly publishes our findings. We are actively seeking talented researchers to join us in shaping this vision. As a Research Scientist focused on Frontier Risk Evaluations, you will design and create evaluation measures, harnesses and datasets for measuring the risks posed by frontier AI systems. For example, you might do any or all of the following: Design and build harnesses to test AI models and systems (including agents) for dangerous capabilities such as security vulnerability exploitation, CBRN uplift, and other high-risk activities; Work with government agencies or other labs to collectively scope and design evaluations to measure and mitigate risks posed by advanced AI systems; Publish evaluation methodologies and write technical reports for policymakers. Ideally you’d have: Commitment to our mission of promoting safe, secure, and trustworthy AI deployments in the industry as frontier AI capabilities continue to advance. Practical experience conducting technical research collaboratively. You should be comfortable building and instrumenting ML pipelines, writing evaluation harnesses, and quickly turning new ideas from the research literature into working prototypes. A track record of published research in m
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a Sr. Staff Design Verification Engineer to lead end-to-end verification efforts for our advanced CPU and AI compute platforms. This role is ideal for seasoned engineers with deep expertise in CPU verification who excel at driving complex test plans to closure, mentoring teams, and leveraging modern AI tools to accelerate innovation. This role is hybrid, based out of Boston, Toronto, Ottawa or Santa Clara. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proven track record leading end-to-end CPU verification across processor cores, caches, interconnects, interrupts, MMU/IOMMU, and cache coherency protocols, with multi-chip and emulation experience a plus. Expert in creating comprehensive test plans and building simulation testbenches using SV-UVM, C/DPI, and Cocotb. Skilled in managing high-volume regression execution, complex failure triage, and driving test plans and coverage metrics to closure. Experienced leader and mentor dedicated to guiding junior engineers and fostering verification best practices across teams. Proficient in leveraging modern AI tools like Copilot, Cursor, Claude, Gemini, and ChatGPT to optimize ver
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to validate the System Management Controller (SMC) and enable seamless multi-chip integration. In this role, you will design and execute tests, build infrastructure, and debug issues across chiplet-based SoCs. You’ll have the opportunity to work with remote mentorship while contributing to the foundation of scalable multi-die systems. This role is hybrid, based out of Toronto, Ontario, Boston, MA or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proficient in SystemVerilog, SV-UVM, Python, and C/C++ with strong verification skills. Experienced in writing test plans, building infrastructure, and debugging hardware/software flows. Comfortable working with remote mentorship and distributed teams. Familiar with AI-assisted tools like Copilot, Cursor, and Claude to accelerate verification. What We Need Develop and maintain SMC tests and supporting DV infrastructure. Write, execute, and track test plans for chiplet and multi-chip SoC designs. Use C/C++ to develop tests compiled, loaded, and executed directly on the DUT. Triage, analyze, and debug issues in clos
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a person ready to take up the challenge of working in a high-profile project where we integrate multiple chiplets into a System-in-package, in collaboration with external stakeholders. You will work with Tenstorrent worldwide experts and leaders in the USA, Japan and other countries, and help us make our IP even better. In this role, you will ensure functionality and performance of the system. This role is based out of Tokyo, Japan and offer flexible Hybrid work style. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Responsibilities: Verify Tenstorrent’s digital IP and SoC logic at chiplet integration level, using industry standard verification methodologies Build and improve components of verification infrastructure including model builds and simulation/regression runs, Create verification components like testbenches, checkers and test generators, Add assertions and coverages along associated methodologies, Build verification test plans for subsystems, align them with project stakeholders, implement test suites, summarize the results and share feedback with project stakeholders Publish and review verification metrics and drive
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We’re looking for a hands-on RTL Design Engineer to own the microarchitecture and RTL implementation of the Power Management Subsystem/Interrupt Controllers/AXI Interconnect/Cache Controller . You’ll collaborate with cross-functional teams—architecture, firmware, software, DV, and PD—to define, design, and optimize power management solutions for next-generation RISC-V/ARM-based SoCs. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels. During the interview process, you will beevaluated and offered a level that aligns with your experience, which may differ from the one in this posting. Who You Are 5–10 years of experience in ASIC/SOC/IP design, with hands-on experience in microarchitecture and RTL development Skilled in Verilog/SystemVerilog and comfortable working across design, debug, and analysis Design and develop microarchitectures for a set of highly configurable IPs Microarchitecture and RTL coding ensuring optimal performance, power, area Work with verification teams on assertions, test plans, debug, coverage, etc. Deep understanding of power management concepts—clocking, reset, DVFS, and low-power modes Familiar with RISC-V or ARM-based SoCs and standard bus protocols (AXI, AHB, APB, CHI) Awareness of functional safety (ISO 26262) practices in hardware design What We Need Abil
MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn healthtech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really is : We are seeking a detail-oriented and technically strong AI QA Engineer to ensure the quality, reliability, and performance of Large Language Model (LLM)-based systems. In this role, you will be responsible for designing and executing test strategies, validating model outputs, and building evaluation frameworks to enhance the accuracy, safety, and overall performance of AI-driven applications.We would particularly value candidates who have hands-on experience in developing evaluation frameworks (evals) for AI systems, along with strong expertise in comprehensive system testing and quality assurance practices.You are responsible for making MeltPlan work in the real world. What You'll Do: Design, develop, and execute evaluation frameworks (Evals) for Large Language Models (LLMs) and AI syst
About Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed system The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties Develop, own, and maintain tools and services to support the software org
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. The Director, Sales Performance & Enablement will own the strategy and systems that drive sales performance across CLEAR’s airport network, translating commercial priorities into playbooks, tools, incentives, and field execution that improve conversion and sales volume. Reporting to the VP, Sales & Service, you’ll combine data-driven strategy with hands-on field leadership to identify what works, test it quickly, and scale it across the network. What you'll do: Own CLEAR’s airport sales execution strategy, developing scalable playbooks, tools, quality standards, and performance frameworks that drive conversion and sales volume across the network Analyze sales performance, market dynamics, and field insights to set network targets, identify performance opportunities, prioritize investment, and translate data into clear actions for the field Build a rigorous test-and-learn engine across high-volume airports, rapidly piloting sales approaches, scripts, tools, and incentives and using results to determine what scales network-wide Partner with Field Operations leaders to embed sales standards, coaching rhythms, and performance cadences into existing station operations, driving consistent execution through influence and shared accountability Lead a high-performing team focused on sales performance intelligence and field quality, while partnering with People and Finance to design commission structures, incentives, and recognition programs that reinforce the right front-line behaviors How you'll measure success: Network-wi
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio
Get new test job 02 09 jobs by email
Daily job updates · Unsubscribe anytime