Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Get to know the team The Developer Platform team at Auth0 (an Okta company) owns the surfaces developers use to build with Auth0 - APIs, SDKs, CLIs, and the tooling that shapes how developers experience identity. Software is increasingly built by developers working alongside AI agents, and that shift raises the bar on every interface we ship: APIs have to be consumable by agents, tools have to be discoverable, and our platform has to hold up when an agent - not just a person, is the one calling it. We're looking for an engineer who already builds with those skills. We move fast, we own problems end-to-end, and we care deeply about the developer experience we put in front of people. The opportunity We're hiring a Staff Software Engineer to build the platform surfaces that developers and AI agents use to configure, extend, and interact with Auth0. You'll work at the intersection of API platform design, developer tooling, and emerging standards like MCP (Model Context Protocol). You'll take on ambiguous, high-leverage problems alongside a group of strong senior and staff engineers, and you'll have real influence over how our platform is built for a world where agents are first-class consumers. What you'll be doing Design and build developer-facing platform surfaces - APIs, MCP tools, CLI integrations, that hold up whether the caller is a developer or an agent Own technical direction for key platform surfaces that make Auth0 consumable by agents Drive cross-cut
Jobiba hiring network
Inference Technical Lead Jobs
1,448 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current inference technical lead jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta’s Workforce Identity Cloud Security Engineering group is looking for an experienced and passionate software security engineer to join a team focused on designing and developing Security solutions to harden our frameworks & infrastructure. We embrace innovation and pave the way to transform bright ideas into excellent security software solutions that help run large-scale, mission-critical software. We encourage you to prescribe defense-in-depth measures, industry security standards, enforce the principle of least privilege to help take our Security posture to the next level. Our Security engineering team has a niche skill-set that combines Security domain expertise with the ability to design, implement and rollout security features and functionalities without adding friction to product functionality or performance. We are responsible for the ever-growing need to improve our customer safety and privacy by providing security services that are coupled with the core Okta product. This is a high-impact role in a security-centric, fast-paced organization that is poised for massive growth and success. You will act as a liaison between the Security org and the engineering org to build technical leverage and influence the security roadmap and direction. You will focus on engineering security and privacy aspects of the systems used across our services while working on a weekly release cadence. You will be empowered to propose stimulating new
Who We Are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About The Team The Billing Solutions Architects (SA) team is a group of specialists with a deep understanding of Stripe's products and the broader subscription, billing, and payments industry. They excel at using Stripe's platform to solve complex customer problems and collaborate with sales teams to create custom solutions and provide strategic guidance to merchants. Their contributions extend to helping achieve revenue and pipeline goals, offering insights on go-to-market plans, and providing valuable feedback to product and engineering teams. What You'll Do You will be a key technical advisor and solutions expert, driving the adoption of our billing solutions with new and existing customers. You'll collaborate closely with sales, marketing, product, and engineering teams to ensure successful customer go-lives and influence our product roadmap based on market trends and customer needs. This involves building deep relationships with customer stakeholders, developing and delivering technical solutions, and scaling the organization's expertise through training and knowledge sharing. You’ll have the opportunity to shape the future of internet commerce - working on challenging problems at a global scale in a collaborative and innovative work environment. Responsibilities: Technical Expertise: Serve as the subject matter expert on our billing and revenue recognition solutions, demonstrating deep understanding of order-to-cash processes, customer journey
About the Team The Personalization-Memory team, within OpenAI's broader Personal AGI organization, is focused on developing agents that can learn from prior interactions in order to become more helpful and efficient over time. We build general-purpose memory and personalization capabilities that transfer across ChatGPT and other agentic products, and we collaborate with applied engineering on the product surfaces that allow users to interact with memory. About the Role As a Research Engineer / Research Scientist on the Personalization-Memory team, you will research and develop improvements to memory usage and personalization in OpenAI's frontier models. Our team works on reinforcement learning, dataset creation, evaluations, and other post-training methods. We partner closely with research and product teams across the company to realize the vision of a truly personalized ChatGPT. We're looking for individuals who have a background in frontier model post-training, are able to iterate quickly, and who are passionate about product-driven research. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda for improving memory use and personalization in frontier models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the research and product teams to influence the shape of technical solutions in the product. You might thrive in this role if you: Are passionate about personalization and building personalized assistants. Have experience working with user signals and human data to turn feedback into reliable signals for training and evaluation. Have a deep understanding of frontier model post-training and machine learning applications. Value principled approaches and research craftsmanship. Are comfortable diving into a lar
About the Team The Personalization-Memory team, within OpenAI's broader Personal AGI organization, is focused on developing agents that can learn from prior interactions in order to become more helpful and efficient over time. We build general-purpose memory and personalization capabilities that transfer across ChatGPT and other agentic products, and we collaborate with applied engineering on the product surfaces that allow users to interact with memory. About the Role As a Research Engineer / Research Scientist on the Personalization-Memory team, your work will span memory architecture, post-training, and developing long-horizon tasks for training and evaluations. We're looking for individuals who have a background in reinforcement learning research, are able to iterate quickly, and who can convert scientific rigor and long-term research into realized product impact. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own and pursue a research agenda for improving long-horizon memory and personalization in frontier models. Build robust evaluations for tracking modeling improvements. Design, implement, test, and debug code across our research stack. Collaborate closely with the research and product teams to influence the shape of technical solutions in the product. You might thrive in this role if you: Love being on the cutting edge of RL and frontier model research. Value principled approaches and research craftsmanship. Are passionate about long-horizon tasks, memory, and turning your research into product impact. Are comfortable diving into a large ML codebase to debug. Thrive in a fast-paced, dynamic, and technically complex environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI syst
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Streamlit in Snowflake (SiS) is our flagship app development platform, enabling data engineers, analysts, and developers to build and deploy interactive, data-driven applications directly within Snowflake. As Principal Engineer for the Streamlit in Snowflake team, you will be the top technical voice shaping how millions of users build and deploy apps on Snowflake's data platform. This role requires a rare combination: deep hands-on engineering expertise in Python/container runtimes, a systems architect's instinct for platform design, and the organizational influence to drive cross-team programs. You will define how SiS evolves from its container-native SPCS runtime and embedding/iframe SDK, to its developer experience, performance at scale, and integration with Snowflake's broader AI and data ecosystem. AS THE PRINCIPAL ENGINEER FOR STREAMLIT IN SNOWFLAKE YOU WILL: Define and own the architectural vision for the SiS platform spanning the SPCS container runtime, the warehouse runtime, the embedding SDK, developer tooling, and the Snowsight integration layer. Drive the platform's evolution toward its next-generation capabilities building on publicly shipping features like chromeless viewer URLs and IdP integration, and shaping the architectural direction for areas still in de
Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role As a Wireless Systems Software Engineer , you will help build the core software libraries and system infrastructure that power the next generation of wireless experiences at HP IQ. You will work across hardware, firmware, embedded software, operating systems, and product teams to integrate and optimize multiple wireless technologies, including Bluetooth, Ultra-Wideband (UWB), NFC , and future connectivity solutions. You will play a key role in designing scalable software architectures that bridge hardware and software, enabling seamless communication across embedded devices and host platforms. This position offers the opportunity to work on challenging system-level problems, influence the architecture of a new wireless ecosystem, and help define technologies that will shape the future of work. What You Might Do Serve as a technic al expert across wireless technologies including Bluetooth, UWB, NFC , and related embedded communication interfaces. Architect, design, and develop reusable, scalable wireless software libraries with well-defined API
At Lyft, our purpose is to serve and connect. We aim to achieve this by cultivating a work environment where all team members belong and have the opportunity to thrive. Machine Learning is at the heart of Lyft’s products and decision-making. Machine Learning Engineers at Lyft operate in dynamic environments, moving quickly to build the world’s best transportation solutions. We tackle a wide range of challenges, from pricing and marketplace frameworks that ensure reliability and competitiveness, to agentic AI platforms that automate analytical workflows, to behavioral detection systems that protect the integrity of our network. We operate at the intersection of applied ML and real business impact, shipping models that directly influence revenue, rider experience, and partner trust. Lyft Business builds products that help organizations move the people who matter most—employees, customers, patients, and guests—easily and efficiently. Our offerings include Business Travel, Lyft Pass, and Concierge (for healthcare and non-healthcare rides), enabling companies to manage transportation at scale through APIs, integrations (e.g., Concur, Expensify), and dedicated tools. These platforms power high-impact B2B use cases across corporate travel, healthcare access, customer experience, and community programs. We're looking for a Machine Learning Engineer to design, build, and deploy ML systems across Lyft Business. This is a high-scope role: you won't be siloed into one problem area. Instead, you'll move across pricing algorithms, fraud and behavior detection, agentic AI systems, and emerging ML applications as the business evolves. You'll write production-quality code, own models end-to-end from prototyping through deployment, and collaborate closely with Data Scientists, Product Managers, and Software Engineers to translate complex business problems into scalable ML solutions. This role is ideal for someone who is technically versatile, energized by variety, and wants to see th
About Us Oliv.AI is a SalesTech global startup headquartered in San Francisco, debuting the world's first team of AI Agents for sales. With our recent $5.2M Seed funding, we solve one of the biggest problems for revenue teams: unreliable deal data. Oliv captures Deal Intelligence from every meeting, call, and email—without any rep involvement. The result is a clear, detailed view of every deal, presented in scorecards built on trusted sales methodologies like MEDDICC, BANT, and SPICED. Our AI agents are built for sales teams—sales managers, AEs, and RevOps—handling the work that takes them away from selling. With Oliv AI, sales teams can bring back focus on deals, strategy and conversation. The role This is not a traditional marketing role, and it is not a pure engineering role either. You will decide which accounts matter, build the systems that research them, write the content that reaches them, build the partnerships that amplify it, and turn all of that into qualified pipeline. Two things have to be true about you at once. You are technically sharp enough to build your own workflows, wire up your own APIs, and ship a working system without waiting on engineering or anyone else. Nothing you own should be blocked on a single person, including us. And you are obsessed with distribution — content, syndication, social, and partnerships — because pipeline goes to whoever the buyer already trusts. That means newsletters, communities, podcasts, comparison pages, partner audiences, and the people who influence them. What you will own GTM strategy and experimentation Develop and continuously refine Oliv's ICP, market segments, buyer personas, and account-selection criteria. Translate company goals into specific growth hypotheses and campaign plans. Identify new audiences, buying signals, use cases, and distribution opportunities. Build a structured experimentation roadmap across content, social, syndication, partnerships, outbound, and ABM. Define success metric
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. We are looking for a Senior Product Manager to join our Applied AI organization at Smartsheet. You'll own the roadmap for specialized AI agents and evaluation infrastructure, ship complex agent capabilities end-to-end, and define the standards that enable teams across the company to build agents successfully. You'll report to our Director of Product Management for Applied AI located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Own the strategy, roadmap, and execution for specialized AI agents and evaluation infrastructure Ship large, technically complex agent capabilities end-to-end with with a high degree of autonomy Define agentic development standards that enable teams beyond your own to build high-quality agents Mentor and develop junior PMs on your team Build strong cross-functional partnerships with engineering, design, and data science to drive alignment and move work forward Support other duties as assigned You Have: 8-10+ years of product management experience with technically complex platforms Fluency in AI/ML concepts -model evaluation, agent architectures, RAG, prompt engineering, quality measurement Track record of shipping enterprise software from inception to launch Experience building or managing evaluation and quality infrastructure for AI or ML systems (strongly preferred) Comfort operating in high-ambiguity environments with significant autonomy Exceptional communication skills and the ability to influence without auth
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving team at Cohere. The team is responsible for developing, deploying, and operating the AI platform delivering Cohere's large language models through easy to use API endpoints. In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments. You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs. As a Site Reliability Engineer you will: Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments. Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. Take steps required to ensure we hit defined SLOs, including pa
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads. You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference. Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations. You may
About the Role We are a small team of AI builders in Paytm Labs. As a Staff AI Platform Engineer, you will work across inference and agentic systems. You will contribute to Paytm's AI inference platform (Pi), serving internal teams and enterprise customers - running our own coding and domain-specific models (voice, vision, risk, fintech workflows) as well as third-party models. You will also architect and build the platform that enables autonomous AI agents to operate safely and reliably in production - the runtime, orchestration, and developer tooling for agents to reason, plan, use tools, and execute complex multi-step workflows, automating both software development and business processes. You will work at the intersection of LLMs, distributed systems, and production fintech infrastructure, helping define how inference and agentic AI are built and deployed across payments, risk, fraud, collections, support, and developer experience.
About ElevenLabs ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $11B - multiples of 11, always. We have expanded from voice into three main platforms: ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale. ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages. ElevenAPI gives developers access to our leading AI audio foundational models. Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you. How we work High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy. Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you. AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. Excellence everywhere: Everything we do should match the quality of our AI models. Global team: We prioritize your talent, not your location. What we offer Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible. Growth paths: Joining ElevenLabs means joining a
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Product Engineer on the Dedicated Inference team, you'll shape the state-of-the-art developer experience for deploying and operating AI workloads in production. From the CLI and SDKs to APIs, observability, and debugging workflows, you'll build the tools customers rely on every day to manage mission-critical inference deployments. Few teams at Baseten have as much breadth and visibility as Dedicated Inference. The team is often at the forefront of new product development, giving engineers the opportunity to shape the experience of some of our most important customers. EXAMPLE INITIATIVES You'll get to work on these types of projects as part of our Dedicated Inference team: Chains for multi-component workflows Asynchronous inference Model APIs for frontier models Model training built for production inference RESPONSIBILITIES Implement new features and products for the team Design ergonomic APIs and abstractions to solve customer problems Fix bugs and resolve customer issues with urgency Work across the stack - regardless of where you start, you’ll end up touching both React Components and Kubernetes Pods Work closely with the product and forward deployed engineering teams to develop and drive new product ideas REQUIREMENTS Bachelor's degree or higher in Computer Science or related field Proficient coding abilities in one or more popular programming or scripting languages; Python, Go, or Javascript proficie
Get new inference technical lead jobs by email
Daily job updates · Unsubscribe anytime