About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve. We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure. This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure. responsible for delivering the software but also for operating and supporting it in production. Why this Role You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective. You’ll have the opportunity to build foundational systems from the ground up, work on infrastructur
Jobs in India
The Cape Surface in India
3,773 active opportunities · Updated October 2026
Showing
15 jobs
Explore current the cape surface jobs across India. Filter by work mode, employment type, experience, department, date posted and distance.
About the Role REMOTE IN INDIA We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency. The core value we're after is decoupling the people building on top of the platform from the operational and runtime complexity underneath, while squeezing more usable capacity out of the same hardware. You'll build the controllers, reconciliation loops, and self-service surface (API/CLI, not tickets) that make that decoupling real, plus the event-driven health, remediation, and utilization systems that keep it running and efficient without a human in the loop. Strong candidates have hands-on experience with Kubernetes controller/CRD patterns, have built or operated a platform API that abstracts multiple backends behind one interface, understand GPU scheduling and capacity efficiency (fragmentation, bin-packing, right-sizing), and think about GPU infrastructure as software to be engineered. A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship. You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production. Responsibilities Build the provisioning state machine
About the Role At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale. If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk. Responsibilities Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention. Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation. Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers. Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators. Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads. Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations. Continuously improve deployment velocity, reliability, and operational efficiency through automation. Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure. Requirements 3+ years building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust . Experience building platforms, automation systems, or developer infrastructure. Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies. Strong systems thinking with the ability to understand problems across hardw
About the Role We are looking for a skilled and detail-oriented Python Developer to join our team. In this role, you will focus on writing efficient scripts, analyzing complex data, and delivering actionable insights through clear and accurate visualizations. The ideal candidate will be a fast learner with a strong foundation in software development principles and a passion for solving challenging problems. Key Responsibilities Design and develop efficient Python scripts for data analysis and automation. Build accurate and insightful data visualizations to support decision-making. Implement and maintain alerting tools to monitor data workflows and systems. Collaborate with cross-functional teams to gather requirements and deliver solutions. Apply strong knowledge of data structures, algorithms, and OOP concepts to write performant, maintainable code. Continuously optimize and improve data processing pipelines for scalability and performance. Key Requirements Experience: 2–5 years as a Software Engineer or Python Developer. Strong proficiency in Python with the ability to quickly learn new frameworks and libraries. Hands-on experience with SQL for querying and data manipulation. Experience with data visualization/plotting tools (e.g., Matplotlib, Seaborn, Plotly). Familiarity with alerting and monitoring tools . Solid understanding of data structures, algorithms, and OOP concepts . Excellent problem-solving skills, analytical thinking, and attention to detail. Benefits: Our open and casual work culture gives you the space to innovate and deliver. Our cubicle free offices, disdain for bureaucracy and insistence to hire the very best creates a melting pot for great ideas and technology innovations. Everyone on the team is approachable, there is nothing better than working with friends! Our perks have you covered. Competitive compensation 6 weeks of paid vacation Monthly after work parties Catered breakfast and lunch Fully stocked kitchen International team outing
A Career with Point72's Technology Team As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. What you'll do Lead the design, development, and operation of scalable, enterprise-grade AI/ML architectures and systems with a strong emphasis on reliability, availability, and performance. Lead and mentor a team of engineers, driving technical direction, code quality, and iterative delivery of large-scale solutions. Partner closely with data scientists, engineers, product teams, and compliance to integrate AI/ML solutions into existing and new products. Own the end-to-end lifecycle of GenAI services, including LLM inference, model serving, and proxy/gateway layers that support multiple downstream applications. Define and uphold engineering best practices around observability, scalability, security, and cost efficiency for AI/ML platforms. Evaluate tools, technologies, and processes to ensure the highest quality and performance of AI/ML systems. Stay abreast of the latest advancements in AI/ML technologies and methodologies, and translate them into pragmatic solutions for the business. Ensure compliance with industry standards and best practices in AI/ML. What's required Bachelor's or Master's degree in Computer Science, Engineering, or a related field. 10+ years of experience in software/AI/ML engineering, with a proven track record of successful delivery of complex, production-grade systems. Demonstrated experience building large-scale enterprise-grade services with high reliability, availability, and observability (SLO/SLA-driven en
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. At Tenstorrent, we build open, state of the art compute for real workloads and real developers. You will lead our Bangalore-based CPU Core Design Verification team, driving both the technical strategy and the people leadership behind verification of our high-performance, out-of-order RISC-V CPUs. You will manage a team of core-level verification engineers, partner closely with our Santa Clara and Austin sites, and own delivery of core verification from architecture definition through tapeout readiness. This role is hybrid based out of Bangalore, India, with regular collaboration across our globally distributed CPU organization. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are 15+ years of CPU/DV experience, including 7+ years leading verification teams, with strong hands-on expertise in high-performance, out-of-order CPU verification. Deep microarchitecture & verification expertise, including ISA, RTL, UVM/CVM, testbench development, regressions, coverage, and complex debug. Strong people and technical leadership, with experience hiring, mentoring, developing engineers, and setting verification strategy and methodology across teams. Highly
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are seeking an experienced and motivated System IP RTL Design Lead to drive the definition and implementation of complex datacenter class IP for our next-generation semiconductor products. The ideal candidate possesses a deep understanding of System IP and SoC architecture, IP integration challenges, and is adept at leading a design team to deliver high-quality, scalable IP solutions. This role requires strong hands-on technical leadership, excellent communication skills, and a proven track record in the successful delivery of silicon IP. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A seasoned IP / SoC design leader with 20+ yrs of strong expertise in complex SoC architectures, high-speed protocols, and large-scale IP integration. Skilled in RTL design and debug (Verilog/SystemVerilog) with hands-on experience using industry-standard tools for simulation, synthesis, lint, CDC, and power analysis on high-frequency designs. Experienced in driving microarchitecture decisions, authoring design specifications, and performing complex PPA trade-off optimizations. A proactive problem so
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We’re looking for a hands-on RTL Design Engineer to own the microarchitecture and RTL implementation of the Power Management Subsystem/Interrupt Controllers/AXI Interconnect/Cache Controller . You’ll collaborate with cross-functional teams—architecture, firmware, software, DV, and PD—to define, design, and optimize power management solutions for next-generation RISC-V/ARM-based SoCs. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels. During the interview process, you will beevaluated and offered a level that aligns with your experience, which may differ from the one in this posting. Who You Are 5–10 years of experience in ASIC/SOC/IP design, with hands-on experience in microarchitecture and RTL development Skilled in Verilog/SystemVerilog and comfortable working across design, debug, and analysis Design and develop microarchitectures for a set of highly configurable IPs Microarchitecture and RTL coding ensuring optimal performance, power, area Work with verification teams on assertions, test plans, debug, coverage, etc. Deep understanding of power management concepts—clocking, reset, DVFS, and low-power modes Familiar with RISC-V or ARM-based SoCs and standard bus protocols (AXI, AHB, APB, CHI) Awareness of functional safety (ISO 26262) practices in hardware design What We Need Abil
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. At Tenstorrent, we build open, state of the art compute for real workloads and real developers.You will own CPU core‑level verification, shaping how our out‑of‑order RISC‑V CPUs behave in silicon. This role is hybrid, based out of Bangalore, India. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You bring 8+ years in CPU verification or closely related digital design. You know high‑performance out‑of‑order CPU microarchitecture in depth. You work comfortably with RTL, waveforms, logs, and complex debug scenarios. You communicate clearly across design, DV, emulation, and post‑silicon teams. What We Need Plan and drive functional verification for CPU core features and complex microarchitectural scenarios. Develop UVM, assembly, and C/C++ based stimulus, functional models, and coverage for ISA, RISC-V extensions, and un-core components. Debug simulation and emulation regressions using RTL understanding, waveforms, and logs to identify and resolve issues efficiently. Build and enhance coverage models, testbenches, and debug infrastructure to improve verification quality and coverage closure. Collaborate with design, validation, and
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building the future of AI compute, and this role will build the talent pipeline that makes it possible. As an Early Talent Recruiter, you will own university and early-career recruiting for your site, shaping which schools we invest in, building relationships with students and faculty, and leading intern, co-op, and new-grad hiring from first outreach through accepted offer. You will not simply run an existing campus calendar. You will decide where to focus, raise the quality of the funnel, and build the recruiting motion that our growing hardware and software teams need. This role is hybrid, based out of Tokyo, Japan, Bangalore, India, or New Taipei City, Taiwan. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You have owned a full campus recruiting cycle, including school selection, university partnerships, sourcing, screening, interview coordination, offer management, and closing. You use judgment and influence to partner with hiring managers, challenge assumptions when needed, and hold a clear hiring bar without relying on authority. You work from funnel data and can connect measures such as applications, conversions
About the Role At FourKites we have the opportunity to tackle complex challenges with real-world impacts. Whether it’s medical supplies from Cardinal Health or groceries for Walmart, the FourKites platform helps customers operate global supply chains that are efficient, agile and sustainable. Join a team of curious problem solvers that celebrates differences, leads with empathy and values inclusivity. As a Senior Customer Engineer, you own the technical customer relationship end to end. You run discovery independently, design integration and agentic workflow architectures for complex enterprise problems, and deploy solutions live with customers — often before they know exactly how to articulate what they need. You write optimized, production-grade code at speed, grounded in strong data structures and algorithms fundamentals, because compressing the time from customer problem to working solution is how FDE delivers its value. You bring genuine innovation to hard problems — your solutions are technically sound, elegant, and often non-obvious. You are the primary technical contact for 2–4 major enterprise accounts, you mentor FDEs on the team, and you are building the skills that will take you into Staff-level technical leadership. What You'll Do Own complex integration and AI agent implementations end-to-end — from technical discovery through go-live and post-launch enhancement — as the primary technical contact for 2–4 enterprise accounts Design integration architectures with explicit attention to error handling, retry logic, observability, failure recovery, and multi-system authentication Design and deploy agentic AI workflows that orchestrate supply chain operations — from requirements through production, including regression testing and validation before each customer deployment Deploy AI agent workflows live with customers present — configuring and troubleshooting in the room during customer calls, not gathering requirements to build later Run customer discovery
ABOUT THE UI DESIGNER ROLE: WIN is looking for a talented UI Designers to create modern, intuitive, and visually compelling digital experiences across our web and mobile products. You will work closely with Marketing, UX, Product, and Engineering teams to transform concepts, user flows, and product requirements into polished, functional interfaces. We are looking for someone who combines strong visual design fundamentals with curiosity and speed. You should enjoy experimenting with new ideas, embracing emerging design technologies, and using AI-powered tools to explore design directions, generate references, rapidly prototype concepts, and iterate faster. As we continue to build AI-enabled products and experiences, you will play an important role in defining how those experiences look, feel, and come to life for our users. KEY RESPONSIBILITIES: Design intuitive, visually compelling, and high-quality user interfaces for web and mobile products. Translate UX wireframes, user flows, product requirements, and concepts into polished UI designs and interactive prototypes. Create UI mockups, layouts, components, visual concepts, and prototypes that clearly communicate the intended product experience. Use AI-powered tools to generate design references, explore multiple creative directions, rapidly prototype ideas, and accelerate design iterations. Collaborate closely with UX Designers, Product Managers, and Engineers to take ideas from concept to production. Build and evolve reusable components, design systems, visual guidelines, and UI patterns to ensure consistency across products. Apply strong visual design principles across typography, hierarchy, spacing, color, imagery, and interaction states. Iterate rapidly based on product feedback, user insights, data, and technical considerations. Stay curious about emerging UI trends, AI tools, design technologies, and interaction patterns, and identify opportunities to incorporate them into our products. Maintain high standards
ABOUT THE INSTRUCTIONAL DESIGNER ROLE: WIN is looking for experienced Instructional Designers to create engaging, practical, and impactful learning experiences for our franchise owners, external partners, team members, customers, and other stakeholders to support our growth in the U.S. and internationally. You will work closely with subject matter experts and cross-functional teams to transform complex business processes, product knowledge, and operational concepts into engaging and effective learning experiences. From e-learning modules and videos to instructor-led sessions, onboarding programs, job aids, assessments, and product training, you will help determine the right learning approach for different audiences and needs. We are looking for team members who are curious, creative, and passionate about how peoplelearn. You should be excited to experiment, leverage AI, and continuously improve learning content faster and more effectively. KEY RESPONSIBILITIES: Design and develop engaging learning experiences including e-learning modules, instructor-led training, videos, onboarding programs, product training, job aids, SOPs, assessments, quizzes, and other learning resources. Manage learning content from concept through development, review, launch, and continuous improvement. Create learning journeys, course outlines, storyboards, scripts, activities, simulations, and assessments aligned with specific learning objectives. Utilize innovation and AI tools to accelerate research, content development, storyboarding, visual creation, video and voice production, assessments, and rapid prototyping of learning experiences. Design learning experiences for U.S. audiences with strong attention to language, context, examples, and cultural relevance. Collaborate with Product, Operations, Marketing, Technology, and other teams to develop training for new products, processes, systems, and initiatives. Maintain and update existing training programs and learning resources as product
ABOUT THE PRODUCT SUPPORT SPECIALIST ROLE: WIN is looking for Product Support Specialists to provide exceptional support to users of our cloud-based applications and digital products. You will develop a deep understanding of our products, business processes, and user needs to help users navigate features, troubleshoot issues, configure applications, and get the most value from our technology. We are looking for people who are curious, resourceful, and highly user-focused. You should enjoy solving problems, learning how products work, and explaining solutions in a simple and empathetic way. You will also be expected to embrace AI and modern productivity tools to research issues faster, improve documentation, identify patterns, and make support more efficient and scalable. KEY RESPONSIBILITIES: Develop deep product and business knowledge to become a Subject Matter Expert across WIN's cloud-based applications, websites, and internal platforms. Support users by understanding their needs, answering product questions, explaining features, and helping them effectively navigate and utilize applications. Troubleshoot product and application issues, identify root causes, and provide timely and practical resolutions. Configure cloud-based applications and product settings based on user and business requirements. Manage and prioritize requests through the ticketing system, ensuring timely follow-up, accurate resolution, and a positive user experience. Use AI and modern productivity tools to accelerate issue research, summarize information, troubleshoot efficiently, and improve the quality and speed of support. Identify recurring questions, issues, and user pain points and use these insights to recommend improvements to products, processes, and support resources. Create and maintain clear product guides, FAQs, troubleshooting instructions, knowledge-base articles, and issue-resolution documentation. Leverage AI to help organize and improve knowledge content, while ensuring infor
Overview The Insights team delivers high-quality research, market intelligence, and data-driven analysis to support clients’ strategic and investment decisions. Working across industries-including healthcare and medtech-the team combines primary and secondary research with deep analytical capabilities to generate actionable insights. Team members collaborate closely with global stakeholders to translate complex information into clear, impactful outputs. The Project Manager – Insights is responsible for overseeing the end-to-end delivery of research projects, ensuring timelines, quality standards, and client expectations are consistently met. This role coordinates across internal teams and stakeholders to manage multiple concurrent projects, streamline workflows, and drive execution. The ideal candidate combines strong organizational and communication skills with an ability to translate complex project requirements into efficient delivery. This is a hybrid position out of our Mumbai OR Pune office What You’ll Do : Manage end-to-end delivery of research and insights projects, ensuring timelines and quality standards are met Coordinate across analysts, researchers, and stakeholders to execute project plans Define project scope, objectives, and deliverables in collaboration with internal teams and clients Track project progress, manage risks, and proactively address bottlenecks Oversee preparation and delivery of client-ready reports, presentations, and outputs Ensure consistency, accuracy, and quality of insights and deliverables Optimize workflows, processes, and tools to improve efficiency and scalability Support resource allocation and prioritize workload across multiple concurrent projects Maintain clear and consistent communication with stakeholders on timelines and outcomes What You Have : Bachelor’s degree from an accredited college/university (e.g., Business, Economics, Life Sciences, or related field) +4 years of experience in project management, consulting, r
Other cities to consider
More places hiring for this role
Get new the cape surface jobs in India by email
Daily job updates · Unsubscribe anytime