NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect -- for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author
Jobs in United States
Inference Engineering And Product Lead in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current inference engineering and product lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid Protect is a real-time fraud intelligence product built on a unique advantage: Plaid’s network-level visibility across bank accounts, devices, identities, sessions, institutions, applications, and financial behavior. Protect helps customers detect first-party fraud, synthetic identities, account takeovers, and coordinated attacks that are difficult to see from a single application, account, or transaction. Trust Index turns that fraud intelligence into real-time fraud scores and actionable attributes. This team builds the systems that make this intelligence possible: low-latency inference, new data and model integrations, customer-facing APIs and attributes, safe rollouts, and feedback loops. Ti3 expanded Plaid’s fraud graph nearly 10x and, in early testing, detected up to 41% more fraud at the same false-positive rate. Learn more about Ti2 and Ti3 . We are a small, high-agency team working closely with Product, Data Science, and Machine Learning. We value demos over docs, conviction over consensus/alignment, builder schedule over meeting-heavy calendars. We’re scrappy and a talent-dense team that has high agency and high ownership. As a Staff Software Engineer on the Protect Core team, you wi
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? We're building the foundational infrastructure that will define how the world thinks about and deploys AI, and we want the sharpest, most curious people to help us do it. As a Lead Data Scientist on our Analytics and Data Insights team, you'll tackle problems that don't have textbook answers yet; shaping go-to-market strategy for technology that's still being invented, designing the experiments that prove or kill our biggest bets, and helping enterprises understand what foundational AI actually means for their bottom line. You'll own the full analytical lifecycle, from framing the right questions and building the models, to leading a team that delivers answers leadership can act on. As a Lead Data Scientist, you will: Drive the mission forward. Own the science: design and lead experimentation programs including A/B tests, multi-armed bandits, causal inference studies, that directly map to product and go-to-market decisions. Build predictive models that matter: develop and deploy models for forecasting, segmentation, propensity scoring, and opportunity sizing across Cohere's core business lines. Lead and grow a tea
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: We are hiring a senior/staff software engineer to help design and build core components of our next-generation knowledge retrieval system built for the AI era – search and retrieval infrastructure that powers high-quality, scalable, and enterprise-grade agentic systems. You’ll build the framework that allows our customers to connect knowledge–synthesized from structured and unstructured data–to modern LLM-powered applications, leveraging the world’s best-in-class vector DB supporting semantic search and hybrid retrieval. This role is ideal for someone who loves backend system architecture, distributed systems, and applied AI infrastructure. It is a high impact role with significant ownership across architecture, performance, and system reliability. Responsibilities: Design and build scalable platform components leveraging advanced retrieval via query planning, semantic and hybrid search, metadata-aware search, and LLM generation Design and build optimized indexing pipelines for structured and unstructured data Build backend services for semantic and hybrid retrieval, knowledge graph construction, and retrieval orchestration Improve retrieval quality through evaluation and observability frameworks Design APIs for internal and external user and agentic consumers Optimize latency, throughput and cost across large-scale inference and retrieval workloads Drive technical direction for reliability and security What You’ll Bring to the Table: To thrive in this role, you don't need to check every single box, but you should be deep
From $180K/yr
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Airbnb, a leader in travel and hospitality, is on a mission to create a world where anyone can belong anywhere. Long-Term Stay & Luxe is being built with 0-to-1 intensity inside a marketplace that already moves at massive scale, and this team needs people who can operate like founders while still respecting the rigor of a mature, two-sided marketplace. The Difference You Will Make: Airbnb is seeking a Staff, Advanced Analytics to be the analytics owner and strategic partner for Long-Term Stay & Luxe — setting the agenda for what gets measured and tested, building the quantitative models that underpin how the business grows, and operating as a peer to product and business leadership. A Typical Day: Own the end-to-end analytics strategy and roadmap for Long-Term Stay & Luxe — decide what to measure, what to test, and where the biggest leverage is. Prioritization decisions should exist because of your recommendations. Build and evolve the quantitative models that run the business: supply-demand balance, marketplace liquidity, pricing and take-rate models, and LTV/forecasting — going beyond dashboards to build the underlying data models and pipelines. Design and analyze experiments across pricing, search and ranking, host onboarding, and guest conversion. Develop causal-inference approaches (holdouts, proxy metrics, quasi-experimental methods) for the many cases where a clean randomized trial isn't possible given lower volumes or longer booking cycles. Act as a proactive strategic partner to product, GM, and cross-functional leaders — surface insights and recomme
About the Team Our economics team is continuously working to improve our understanding of an AI-driven economy. About the Role We are seeking a highly technical Economist to join the OpenAI Economic Research team studying the real-world economic impacts of AI. This role is designed for economists with up to 5 years of professional experience post-Ph.D. who are interested in using novel, large-scale datasets to study how AI is reshaping economic systems. We are looking for candidates with deep expertise in at least one core domain relevant to AI’s economic impact, and an interest in contributing to a broader research agenda spanning labor markets, firm behavior, market dynamics, and macroeconomic change. This is an individual contributor role where the candidate will organize and execute on their own data-oriented projects. You will work at the intersection of economic research, data science, and public policy to produce rigorous empirical work that informs decision-makers across the public, industry, and government. Research Areas of Interest We are particularly interested in candidates with demonstrated expertise in one or more of the following areas: Economic Measurement of AI Impact (e.g., adoption trajectories, labor market transitions, productivity growth, and forecasting/scenario modeling for AI-driven economic change) Macroeconomic Implications of AI (e.g., productivity, technology diffusion, economic growth) AI and the Labor Market (e.g., employment, wages, job search, task-level impacts, skill acquisition) Applicants are not expected to have experience across all domains. We aim to build a team with complementary strengths across these areas. In this role, you will: Design and execute empirical research using large-scale observational or experimental data. Apply causal inference and/or structural modeling techniques to study AI-driven economic change. Collaborate with cross-functional teams to translate research questions into testable frameworks and applic
About the Team The Foundations Research team works on high-risk, high-reward ideas that could shape the next decade of AI. Our goal is to advance the science and data that enable our training and scaling efforts, with a particular focus on future frontier models. Pushing the boundaries of data, scaling laws, optimization techniques, model architectures, and efficiency improvements to propel our science. The Search team sits within Foundations, building agentic search by co-designing model–system interfaces with the core search stack (serving, indexing, retrieval) to translate model intent into reliable, real-world actions. Operating at the frontier of AI and information retrieval, the team develops large-scale systems that transform and index vast corpora, enabling models to reason over global knowledge and act dependably. In close partnership with researchers, we rapidly bring modeling breakthroughs into production and redefine how intelligent systems discover, retrieve, and synthesize information at planetary scale. About the Role We’re looking for a Software Engineer focused on building and scaling retrieval systems. You’ll work with a team of researchers and engineers to develop infrastructure that enables models to retrieve and act on the right information at the right time. This includes designing and operating indexing systems, retrieval pipelines, and serving layers. This work supports retrieval across OpenAI products and research, with direct impact on system performance, reliability, and scale. Responsibilities Build and scale retrieval infrastructure across indexing, serving, and query execution. Develop low-latency, high-throughput systems for real-time model interaction. Partner with research to productionize embedding and retrieval techniques. Support dense, sparse, and hybrid retrieval pipelines. Own system performance, reliability, and observability at scale. Collaborate across Pretraining, Inference, and Product teams to integrate retrieval end-to-e
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is transforming how the world uses data and AI — and the networking and traffic infrastructure that powers these experiences is mission-critical. As a Product Manager focused on Traffic & Networking, you will define how Snowflake delivers secure, reliable, and high-performance connectivity at global scale, including the networking foundations required to support AI-driven products and workloads. You will own the product vision and roadmap for internal traffic management, service-to-service networking, customer connectivity, and performance optimization across multi-cloud environments. A core part of this role is defining and evolving Snowflake’s network strategy to support AI products , including latency-sensitive inference, large-scale model training pipelines, vector search, streaming ingestion, and cross-region data movement. This is a high-impact role at the intersection of distributed systems, cloud networking, and AI infrastructure. AS A PRODUCT MANAGER AT SNOWFLAKE YOU WILL: Define the networking strategy required to support Snowflake’s AI products , including low-latency inference paths, high-throughput data pipelines, GPU-adjacent services, and elastic scaling for AI workloads. Partner with AI platform, compute, and storage teams to ensure networking
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role OpenAI's Hardware organization builds supercompute platforms from silicon and boards to full rack-scale systems to power advanced AI workloads. This role owns end-to-end quality for high-speed interconnect hardware across the product lifecycle: early design influence, supplier/contract manufacturer readiness, qualification, ramp, and fleet quality in lab and data center environments. You will be the quality lead for advanced interconnect components and assemblies, including high-speed copper cables, cable cartridges, patch panels, backplane/cable-backplane solutions, high-speed connectors, and related electro-mechanical interfaces. You will partner closely with electrical, mechanical, SI/PI, systems, reliability, operations, and external vendors to prevent escapes and drive rapid, data-driven containment and corrective action. In this role you will: Own quality for advanced interconnect components and assemblies: high-speed connectors, high-speed copper cables, cable cartridges (e.g., cable cassette style assemblies), patch panels & optics, and backplane/cable-backplane interconnect solutions. Drive quality-by-design: participate in design reviews, DFM/DFx, tolerance stacks, material and plating selections, connector mating strategy, strain relief, and assembly methods to reduce variation and field failures. Define and track quality and reliability metrics (DPPM, yield, escapes, RMA/FRACAS trends, Cpk/Ppk where applicable) for interconnects across NPI and m
Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. The Advanced Packaging Technology Development (APTD) team at Micron is shaping the future of memory and storage technologies that fuel the AI era. We collaborate with global R&D groups, suppliers, and manufacturing teams to turn bold ideas into production-ready solutions. We innovate fast, learn continuously, and work as one team! This position is intended to be part of Micron's Technical Leadership Program (TLP). This is a career path for individuals seeking to advance as technical leaders and industry innovators. TLP members are expected to influence, lead, and mentor others. Position Overview: We are seeking an experienced organizational leader who will l ead and influence methodology improvements, including best practices in process development, data-driven decision-making, and application of AI/advanced analytics to drive innovation. In this role, you will drive strategic process development for Micron’s most advanced memory products, including HBM, 3DS DRAM, and emerging 3D architectures. You will guide end‑to‑end technology execution, partner across organizations, and influence both internal roadmaps and the external ecosystem. This is a high-impact role for someone who thrives at the intersection of innovation, complexity, and collaboration. Responsibilities: Own advanced packaging process technologies across wafer‑level and die‑level modules from pathfinding through manufacturing readiness Develop
Senior Designer Description - ABOUT THE ROLE We’re looking for a Senior Designer who leads by doing. This is an equal-parts strategic and hands-on role — someone who sets the creative bar for our in-house agency, earns credibility through exceptional work, and isn’t afraid to roll up their sleeves when it matters. You champion design excellence, can defend your decisions with clarity and confidence, and help grow the people around you. This role sits within a balanced creative environment: you’ll have real influence over the work, while partnering closely with internal clients across marketing, sales, and product who hold final approval. Success here means earning trust through quality and communication, not just creative instinct. Day-to-day you’ll operate across mid-to-lower funnel channels — social, ecommerce, retail, sales enablement, and short-form video. You think like a creative director but know that craft lives in execution. Production is where the rubber meets the road, and you’re proud to be there. WHAT YOU’LL DO Produce exemplary design work across social, etail, retail, sales decks, and short-form video that raises the quality bar for the whole team Partner with internal clients across marketing, sales, and product — translating briefs into design that performs and building lasting working relationships Concept and execute with both visuals and words — thinking across design and messaging together, not in isolation Defend design decisions with strategic rationale, balancing business goals with creative integrity Contribute upstream — shaping briefs, influencing strategy, and catching creative problems before execution begins Maintain and evolve brand standards, ensuring visual consistency across all channels and touchpoints Build scalable systems — templates, as
Become a part of our caring community The Senior Compliance Professional ensures compliance with governmental requirements. The Senior Compliance Professional work assignments involve moderately complex to complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors. Regulatory Compliance – State Enterprise Intake and Implementation – Senior Compliance Professional The Senior Compliance Professional implements new statues, regulations, rules and other regulatory guidance issued by state and federal regulators across the enterprise. Coordinates business partner engagement and implementation of rules. Researches compliance issues and recommends changes that ensure compliance with regulatory obligations. Provides compliance guidance and direction to business partners. Monitors metrics and other oversight tools that track implementation activity. Recommends new measures of compliance performance. Begins to influence department's strategy. Makes decisions on moderately complex to complex issues regarding implementation components. Exercises considerable latitude in determining objectives and approaches to assignments. The Senior Compliance Professional's primary focus will be to implement new federal and state rules, including PBM, across the enterprise. Key responsibilities may include: Serve as the subject matter expert and point-of-contact for individual state and federal implementations, leading implementation activity from onset to conclusion. Research, understand, and apply laws, regulations, and regulatory guidance for federal and state compliance issues. Analyze business requirements and complex issues, conduct research, and provide regulatory guidance to business partners, Law, Risk, and Compliance associate and leaders with regard to federal
Citi Commercial Bank - Mid-Corp Relationship Manager, Healthcare, Higher Ed & Nonprofit - Senior Vice President
CitiThe Mid-Corp Relationship Manager is a strategic professional who closely follows latest trends in own field and adapts them for application within own job and the business. Typically a small number of people within the business that provide the same level of expertise. Excellent communication skills required in order to negotiate internally, often at a senior level. Developed communication and diplomacy skills are required in order to guide, influence and convince others, in particular colleagues in other areas and occasional external customers. Accountable for significant direct business results or authoritative advice regarding the operations of the business. Necessitates a degree of responsibility over technical strategy. Primarily affects a sub-function. Responsible for handling staff management issues, including resource management and allocation of work within the team/project. Responsibilities: Calls on clients to deepen relationships and proactively owns, responds to, uncovers and anticipates future needs, roadblocks or risks and expectations Introduces solutions to clients in building and strengthening an effective portfolio; Works with product specialists and subject matter experts to structure innovative and customized solutions that meet clients’ individual needs Works closely with Case Manager on the on-boarding and retention of clients, ensuring the appropriate “Know Your Client” (KYC) and other compliance deliverables are met; Identifies cross-sell opportunities to deepen and increase share of wallet; Maximizes client experience by proactive sharing markets updates, trend and intelligence Drives innovation in the solutions we provide clients and further developing our business where necessary and appropriate Execution of strategic initiatives launched centrally at all levels (Group, Bank, commercial market and EIB) Networks with clients to identify
Other cities to consider
More places hiring for this role
Get new inference engineering and product lead jobs in United States by email
Daily job updates · Unsubscribe anytime