Title: Senior AI QA Engineer Location: Bengaluru (Bangalore) Opportunity: As a Senior AI QA Engineer for our Precision Patient Care Pipeline , you will go beyond traditional functional testing. You will be responsible for building the framework that ensures our clinical insights are accurate, safe, and reliable. This role requires a unique blend of high-level software testing and data engineering to validate complex, non-deterministic medical outputs using a hybrid of automated grading methodologies . Key Responsibilities: Architect Multi-Layered Validation Frameworks: Design and implement structured testing strategies that combine deterministic checks, semantic similarity metrics, and model-based evaluations. Automated Model Grading: Develop systems to evaluate clinical pipeline outputs for faithfulness, safety, and hallucination detection using various automated scoring techniques (e.g., BERTScore, ROUGE, or custom heuristics). Vibe-Driven Development: Leverage agentic AI tools to rapidly prototype complex test harnesses, "red-team" clinical logic, and build internal validation utilities at high velocity. Data Pipeline Integrity: Execute integration and regression tests for data-heavy backend processes, ensuring medical data remains consistent from ingestion to insight generation. Collaborative Strategy: Work closely with Data Scientists and Product Managers to define "Ground Truth" datasets and clinical evaluation rubrics. Root Cause Analysis: Deep-dive into complex system failures to identify whether issues stem from code logic, data drift, or model behavior. Requirements: 6+ years of technical experience in Quality Assurance, with a strong focus on system architecture and backend data validation. Advanced Python Proficiency: Expert-level skills in Python for building custom test scripts and working within AI/ML ecosystems. AI/ML Validation Experience: Proven experience testing model outputs using diverse metrics (e.g., Semantic Similarity, NLP metrics, an
Jobs in India
Model Behavior Engineer in Bengaluru
161 active opportunities · Updated October 2026
Showing
15 jobs
Explore current model behavior engineer jobs in Bengaluru. Filter by work mode, employment type, experience, department, date posted and distance.
AI Engineer - Enterprise Search Overview We are looking for an experienced Enterprise Search Lead to build and optimize our multi-tenant enterprise search solution. This role focuses on creating scalable systems that integrate seamlessly with customer environments, leveraging cutting-edge AI and ML technologies. You will design and manage the enterprise knowledge graph, implement personalized search experiences, and drive AI- powered innovations to enhance search relevance and ITSM workflows. What You Will Do: Build and structure our enterprise knowledge graph to organize content, people, and activity into meaningful relationships for better search relevance. Develop and refine personalized ranking models that adapt to user behavior and improve search results over time. Design ways to adapt AI language models to each customerʼs data for enhanced accuracy and context. Explore innovative methods to combine LLMs with search engines for answering complex queries Write clean, robust, and maintainable code that integrates smoothly with multi-tenant systems. Collaborate with cross-functional teams to align search capabilities with ITSM workflows. Mentor junior engineers or learn from experienced ones to grow as a technical leader. What You Should Have 3-5 years of experience working on enterprise search products with AI/ML integration. Expertise in multi-tenant systems and securely integrating with external customer systems. Hands-on experience with tools like Elasticsearch, Solr, or similar search platforms. Strong coding skills in Python, Java, or equivalent languages. A passion for solving complex problems with AI and delivering intuitive user experiences. Important notice for candidates: Job scams are on the rise. Please keep these guidelines in mind when applying for any open roles at Atomicwork. Only apply through official Atomicwork channels. We do not use third-party agencies or individuals who ask for payments in exchange for interviews or offer letter
Staff Engineer - Enterprise Search Overview We are looking for an experienced Enterprise Search Lead to build and optimize our multi-tenant enterprise search solution. This role focuses on creating scalable systems that integrate seamlessly with customer environments, leveraging cutting-edge AI and ML technologies. You will design and manage the enterprise knowledge graph, implement personalized search experiences, and drive AI-powered innovations to enhance search relevance and ITSM workflows. What You Will Do Build and structure our enterprise knowledge graph to organize content, people, and activity into meaningful relationships for better search relevance. Develop and refine personalized ranking models that adapt to user behavior and improve search results over time. Design ways to adapt AI language models to each customer’s data for enhanced accuracy and context. Explore innovative methods to combine LLMs with search engines for answering complex queries. Write clean, robust, and maintainable code that integrates smoothly with multi-tenant systems. Collaborate with cross-functional teams to align search capabilities with ITSM workflows. Mentor junior engineers or learn from experienced ones to grow as a technical leader. What You Should Have 5+ years of experience working on enterprise search products with AI/ML integration. Expertise in multi-tenant systems and securely integrating with external customer systems. Hands-on experience with tools like Elasticsearch, Solr, or similar search platforms. Strong coding skills in Python, Java, or equivalent languages. A passion for solving complex problems with AI and delivering intuitive user experiences. Important notice for candidates: Job scams are on the rise. Please keep these guidelines in mind when applying for any open roles at Atomicwork. Only apply through official Atomicwork channels. We do not use third-party agencies or individuals wh
Level Up Your Career with Zynga! At Zynga, we bring people together through the power of play. As a global leader in interactive entertainment and a proud label of Take-Two Interactive, our games have been downloaded over 6 billion times—connecting players in 175+ countries through fun, strategy, and a little friendly competition. From thrilling casino spins to epic strategy battles, mind-bending puzzles, and social word challenges, our diverse game portfolio has something for everyone. Fan-favorites and latest hits include FarmVille™, Words With Friends™, Zynga Poker™, Game of Thrones Slots Casino™, Wizard of Oz Slots™, Hit it Rich! Slots™, Wonka Slots™, Top Eleven™, Toon Blast™, Empires & Puzzles™, Merge Dragons!™, CSR Racing™, Harry Potter: Puzzles & Spells™, Match Factory™, and Color Block Jam™—plus many more! Founded in 2007 and headquartered in California, our teams span North America, Europe, and Asia, working together to craft unforgettable gaming experiences. Whether you're spinning, strategizing, matching, or competing, Zynga is where fun meets innovation—and where you can take your career to the next level. Join us and be part of the play! Position Overview : Zynga’s data science and analytics team uses our unique and expansive data to model and predict user behavior, making our games more personalized and more fun to play! We strive for a better understanding of our players which translates into challenges and features that delight them and increased social engagement within our games. Here’s where you would come in: identify and formalize problems to understand user behavior. Create systems and analyses to derive insights that change the studio’s world view. Be innovative, be creative, use every bit of that key commodity – data. Millions of people play Zynga games every day, so our data is tremendously rich and we have a lot of it! We will rely on you to communicate your findings to your peers – both technical and non-technical. Your solutions t
A CAREER WITH POINT72’S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU’LL DO As Database Support Engineer, you’ll support various critical database platforms across Development, QA, UAT, and Production environments. The role partners closely with application teams, application support, and database engineers and operates within a Follow‑the‑Sun model to ensure availability, performance, and reliability of database services. Key responsibilities include: • Provide operational support for enterprise database platforms in both on-prem private cloud and public cloud • Monitor database health, capacity, performance, and availability, and respond to alerts, diagnose issues, and perform timely remediation • Perform routine maintenance activities (patching, upgrades, housekeeping etc) • Troubleshoot database‑related incidents and collaborate on root cause analysis • Work closely with application owners, application support teams, and DB Engineers • Provide guidance on database best practices and operational standards • Participate in cross‑team problem resolution and continuous improvement initiatives • Contribute to design, implementation and testing of automation and self service capabilities of DB platforms • Drive continuous improvement, identifying opportunities to reduce toil and increase platform efficiency. • Participate in a Follow‑the‑Sun operating model, including shift‑based coverage and handoffs WHAT’S REQUIRED • Bachelor’s degr
About Bolna Bolna is a YC-backed voice AI orchestration platform built for the Indian market—powering multilingual, vernacular voice agents across Hindi, Hinglish, Tamil, and 10+ languages at sub-500ms latency across collections, recruitment, sales, and e-commerce use cases. We are an orchestration layer, not a model company: our moat is outcome-labelled vernacular data, rigorous evaluation infrastructure, and a growing taxonomy of how Indian enterprise voice AI fails in production. Why This Role Exists Product decisions at Bolna increasingly hinge on rigorous, code-mixed-aware data analysis—and not just one kind. On one side, there is model and evaluation rigor: LLM benchmarking for post-call intelligence, ASR/WER evaluation, inter-rater reliability on human-labelled calls, and routing and latency economics. On the other, there is product and growth insight: understanding where self-serve users drop off in their journey, what patterns emerge across lakhs of monthly calls, and which use cases and configurations are actually working. Both currently sit with the Head of Product alongside strategy and roadmap ownership. We need a dedicated analyst to own the execution and recurring cadence across both-freeing product leadership to act on findings rather than produce them. What You’ll Do Model and Evaluation Analysis LLM and model benchmarking: Run structured comparisons across model providers such as Sarvam, DeepSeek, Gemini, and Claude variants for tasks including post-call extraction and LLM-as-judge scoring. Evaluate cost, accuracy, fill rate, and TTR, with particular attention to Hinglish and code-mixed content. Evaluation infrastructure: Build and maintain LLM-as-judge pipelines using tools such as DeepEval, design and track evaluation metrics, and run inter-rater reliability analysis such as Krippendorff’s alpha across human call reviewers. Golden dataset creation: Support the construction of golden datasets for ASR and transcript labelling, including flagging co
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment. Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done. Forward is currently seeking a Senior Backend Software Engineer to work as part of our Platforms team. You will play a critical role in designing, developing, and scaling the core backend services and infrastructure that support our SaaS and on-prem deployments. Your contributions will have a direct impact on the stability, performance, and scalability of our platform, helping to ensure an exceptional experience for our customers. Responsibilities: Platform development: Contribute to the design and development of storage systems, job scheduling systems, data ingestion frameworks, monitoring frameworks etc to ensure high system performance and availability. Feature development: Build and maintain backend frameworks that support essential platform features Scalability & Reliability: Develop scalable, high-performing
Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment. Our customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast-growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security. Backed by world-class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people-centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI-ready operations. Forward is currently seeking experienced Java developers to work as part of our Network team. Responsibilities Help bring the best ideas from the software development world into the networking industry. Contribute to our code base, systems and software architecture as a member of our engineering team. Help create and optimize network device models for different device vendors and protocols. Help create infrastructure needed to configure, collect and test network devices. Work with peers who are experts in Networking, Distributed Systems, Big Data and Search. Requirements 5+ years of work experience in software development 3+ years of work experience with Java BS in Computer Science or related degree Solid software engineering experience with large code bases Basic understanding of networking and TCP/IP. Strong verbal and written communication skills. Nice to haves Working knowledge of how switches, routers, firewalls or load balancers work. Experience working with networking protocols such as BGP/OSPF/IS-IS, IPv4/IPv6, MPLS, VLAN, VXLAN, etc. This position is a re
Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment. Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done. Forward is currently seeking a Java Backend Software Engineer to work as part of our Apps - Server team. The work will involve developing our web server, REST APIs, and product core by writing clean and solid code that interacts with our other services and components. Responsibilities include: Developing new product features that leverage the network model to help users: visualize their network, understand how it behaves, see how it has evolved, answer specific questions, and plan changes Designing the data model for new product features Proposing and implementing REST APIs to support the Forward web application and to publish to customers Constructively reviewing product designs, technical design documents, and code changes Requirements: At least 5+ years of full lifecycle software development experience Expertise in Java (versi
A Career with Point72's Technology Team As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. What you'll do Lead the design, development, and operation of scalable, enterprise-grade AI/ML architectures and systems with a strong emphasis on reliability, availability, and performance. Lead and mentor a team of engineers, driving technical direction, code quality, and iterative delivery of large-scale solutions. Partner closely with data scientists, engineers, product teams, and compliance to integrate AI/ML solutions into existing and new products. Own the end-to-end lifecycle of GenAI services, including LLM inference, model serving, and proxy/gateway layers that support multiple downstream applications. Define and uphold engineering best practices around observability, scalability, security, and cost efficiency for AI/ML platforms. Evaluate tools, technologies, and processes to ensure the highest quality and performance of AI/ML systems. Stay abreast of the latest advancements in AI/ML technologies and methodologies, and translate them into pragmatic solutions for the business. Ensure compliance with industry standards and best practices in AI/ML. What's required Bachelor's or Master's degree in Computer Science, Engineering, or a related field. 10+ years of experience in software/AI/ML engineering, with a proven track record of successful delivery of complex, production-grade systems. Demonstrated experience building large-scale enterprise-grade services with high reliability, availability, and observability (SLO/SLA-driven en
MeltPlan is building the “planning engine” for the $14 Tn construction industry, an AI system designed specifically to optimize decisions before construction begins. While design software optimizes use and aesthetics and construction software optimizes execution and control, MeltPlan is building the missing layer - software that optimizes decisions and tradeoffs upstream, before scope is locked, procurement begins, and change orders become inevitable. MeltPlan’s long-term goal is to help teams make construction “boring” by making planning more intense: surfacing constraints and tradeoffs early, aligning stakeholders before plans are frozen, and reducing the need for late-stage redlines, rework, and change orders. MeltPlan is founded by operators who have built at scale. Kanav previously co-founded Innovaccer, a $3Bn healthtech company focused on making US healthcare more affordable and accessible. He’s now applying that systems-level thinking to construction.He’s joined by Tanmaya Kala, former Project Executive at DPR Construction, who led large commercial, healthcare, and life sciences projects. We combine deep tech scale with real construction execution. What This Role Really is : We are seeking a detail-oriented and technically strong AI QA Engineer to ensure the quality, reliability, and performance of Large Language Model (LLM)-based systems. In this role, you will be responsible for designing and executing test strategies, validating model outputs, and building evaluation frameworks to enhance the accuracy, safety, and overall performance of AI-driven applications.We would particularly value candidates who have hands-on experience in developing evaluation frameworks (evals) for AI systems, along with strong expertise in comprehensive system testing and quality assurance practices.You are responsible for making MeltPlan work in the real world. What You'll Do: Design, develop, and execute evaluation frameworks (Evals) for Large Language Models (LLMs) and AI syst
Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability. Key Responsibilities MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. Job Description/ Responsibilities: Designing, developing and maintaining stable and reliable AI/ML Ops platforms / pipelines Minimum experience of 4-6 Years required in AI ML Ops Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time. Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production. Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow. Security & Compliance: Implementing data security measures, ensuring compliance with data governance
Couchbase, the operational data platform for AI, empowers businesses to succeed by bringing data to life in new ways. Major market-leading companies rely on Couchbase for mission critical operational, analytical, mobile and AI workloads. Built to replace legacy infrastructure and fragmented data services, Couchbase empowers enterprises with a unified platform architected for performance, flexibility and global scale. With Couchbase, organizations bring their data to life, launching game‑changing customer experiences, exploring the limitless potential of AI, and seamlessly extending applications from the cloud to the edge and beyond. Couchbase’s AI‑ready technology and enterprise partnership model eliminate complexity and reduce total cost of ownership, enabling teams to stay agile, innovative and secure. Couchbase believes data should never slow you down, but act as the foundation for your next breakthrough. Discover why Couchbase is trusted to help the world’s biggest players scale, move fast and stay resilient, no matter what’s next on their roadmap. Visit couchbase.com and follow us on LinkedIn and X. Want to be part of our story? Apply today! As a Solutions Architect , you will be the trusted advisor helping our customers unlock the ultimate value of Couchbase. You will work directly with customer engineering teams, blending high-level architectural coaching with hands-on professional service delivery. If you love solving complex, real-world data challenges, designing cutting-edge AI-ready architectures, and guiding enterprise customers from design to production, this role is for you. What You’ll Do (Responsibilities): Lead Customer Engagements: Guide short- to medium-term on-site and remote client engagements. You will lead architecture and use case reviews, sizing, performance tuning, and deployment topology planning to get Couchbase up, running, and optimized. Solve Complex Technical Challenges: Work hand-in-hand with customers t
Other cities to consider
More places hiring for this role
Get new model behavior engineer jobs in Bengaluru, India by email
Daily job updates · Unsubscribe anytime