NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are looking for a motivated Deep Learning engineer to bring advanced communication technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created communication libraries like NCCL, NVSHMEM & technology like GPUDirect -- for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. Communication performance between the GPUs has a direct impact on AI applications. Your work in AI toolkits will make all of those easier for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Improve AI compilers to hide communications or perform automatic fusion. Conduct in-depth AI workload performance characterization on multi-GPU clusters. Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Author
Jobs in United States
Senior Deep Learning Framework Communications Engineer in United States
15 active opportunities · Updated September 2026
Showing
15 jobs
Explore current senior deep learning framework communications engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning. What you will be doing: Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters. Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components. Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware. What we need to see: M.S./Ph.D. degree in CS/CE or equivalent experience. 5+ years of relevant experience. Excellent C/C++ programming and debugging skills. Strong experience with Linux. Expert understanding of computer syst
Are you a person who likes to work in a fast-paced organization? NVIDIA is the world leader in Visual Computing. We are passionate about four markets: Gaming, Automotive, Enterprise Graphics and HPC/Cloud Datacenters; in addition to our traditional OEM business. We are well positioned as the ‘AI Computing Company’, and our GPUs are the brains powering modern Deep Learning software frameworks, accelerated analytics, big data, modern data centers, smart cities, and driving autonomous vehicles. We have some of the most forward-thinking and talented people on the planet working for us. If you're forward-thinking, hardworking, driven and if working with extraordinary people across countries sounds interesting, this job is for you. We are now looking for a Human Resources Business Partner to provide HR support onsite in Santa Clara, CA for our Worldwide Field Organization in a dynamic and collaborative environment. This is a global organization, and we are looking for someone to be passionate about supporting and building strategies to enable NVIDIA to achieve success. You’ll partner with a cross-functional group of subject matter experts to design and execute strategies for how we staff, onboard, develop, motivate, retain and organize work. You will need excellent communication skills, critical thinking and planning ability, and the agility to function in a fast paced and innovative environment. What you'll be doing: This position will be an integral enabler of the mission of our Field organization. In this position you will work with the senior leaders and leadership teams within NVIDIA organizations to develop and execute the HR strategies that champion organizational and people effectiveness. You will think strategically as well as roll up your sleeves and dive deep into practical application. You must understand business priorities and translate them into an HR ag
$155K – $400K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role As a Senior Software Engineer on Sentry’s AI team, you’ll be directly responsible for developing the platform used by our debugging agents. This role is crucial; you will be at the forefront of integrating AI and machine learning into our core products, from issue triage and resolution to predictive analytics for application performance monitoring. Your work will help companies around the globe gain actionable insights into their software, enabling them to build better products, faster. In this role you will Build state-of-the-art agentic AI platforms to triage, debug, and solve real production issues Leverage Sentry’s novel (and massive) dataset of errors, spans, and profiles Own the development of major initiatives in the AI/ML space You'll love this job if you Are driven by impact and enjoy working on high-stakes, high-visibility projects Enjoy building things. You will have the opportunity to join the AI/ML team as one of its foundational members Thrive in cross-functional teams and enjoy building features alongside developers and product teams Qualifications Minimum 5+ years of professional experience with Bachelor’s degree in computer science, machine learning, or a related field Demonstrated expertise building production-grade agentic systems and tools You are comfortable writing production quality code (we use Python and Typescript) Familiarity with deep learning frameworks (we use PyTorch) Familiarity in deploying machine learning models at scale in production environments The base salary range (or hourly wage range, if applicable) that Sentry reasonably expects to pay for this position is $155,000 to
At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 Role Overview: We are seeking a skilled and experienced Senior AI Engineer - Multi-Agent Frameworks to join our AI Platform team. In this role, you will play a pivotal part in building a cutting-edge platform that empowers our users to create and deploy sophisticated intelligent agents, with a key focus on enabling collaborative and multi-agentic behaviors . This is a backend-focused role that requires deep expertise in AI, large language models (LLMs), and orchestration software. Key Responsibilities: Design, develop, and maintain a robust platform to enable users to create and manage AI agents and their interactions. Integrate and work with multiple LLMs, ensuring seamless orchestration and scalability for both individual and coordinated agent operations. Leverage orchestration frameworks like LangGraph and others to build complex workflows and pipelines that support diverse agent functionalities, including frameworks for multi-agent coordination . Develop and implement evaluation frameworks for testing AI agents in challenging and complex scenarios, focusing on individual performance and system-level dynamics. Stay at the forefront of AI advancements, incorporating the latest research and technologies into our platform to enhance agent capabilities and collaboration. Collaborate with cross-functional teams, including product managers, designers, and frontend engineers, to deliver a seamless user experience for building and deploying intelligent systems. Address challenging AI privacy scenarios, ensuring compliance with data protection regulations and best practices within agent-based applications. C
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Data and Analytics team is currently looking for a Senior Data Scientist to join us! You’ll be responsible for laying the foundation for a best-in-class business analytics function. You’ll partner closely with our business stakeholders to ensure that our analytics stack and processes meet the business needs today with an eye towards the future. What you’ll do as a Senior Data Scientist at Vanta: Build and maintain trusted product data assets using dbt, Snowflake, and modern analytics infrastructure Leverage AI-powered analytics tools and data agents (e.g., Snowflake Cortex) to accelerate insight generation, automate repeatable analysis, and scale decision-making Define and evolve measurement frameworks for product health, customer lifecycle, and AI-powered product experiences Partner closely with Product, Engineering, Design, and Customer Success to influence product strategy through data Help define Vanta’s analytics strategy and AI measurement practices as our product and data platform evolve Lead executive analytics reviews, translating complex analyses into clear recommendations that drive company decisions How to be successful in this role: 4+ years of experience working with data as a Data Scientist, Product Analyst, or Analytics Engineer in an applied business setting Strong foundation in SQL, Python (or R), statistics, and machine learning Experience designing and evaluating experiments, predictive models, and other statistical analyses to inform product decisions Experience building scalable data assets, metrics, and analytical frameworks on modern cloud data platforms (e.g., Snowflake, dbt) Deep experience with da
From $168K/yr
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Total Rewards Compensation team serves as a strategic advisor to business leaders, managers, and employees across Airbnb. We are compensation experts with a deep understanding of our stakeholders, problem solvers who use data and insights to drive value, and partners who build the tools, models, and frameworks that support sound compensation decision-making across the company. This role sits within the Compensation function and will work closely with the Technology organization and cross-functional partners including Recruiting, People Analytics, Finance, Legal and Talent. The Difference You Will Make: We are looking for a Technical Compensation Partner to serve as the dedicated compensation partner for the Technology organization, covering Engineering, Infrastructure, Machine Learning, and Data Science. This role will support VP and Director level tech leaders and their Talent Partners as the primary day-to-day compensation resource. The ideal candidate will be a strong business partner and a builder of compensation programs, tools, and data infrastructure, and must be comfortable operating with autonomy in a fast-moving environment. A Typical Day: Serve as the dedicated compensation partner for the Technology organization, supporting VP and Director level leaders, Talent Directors, and People Partners across Engineering, Infrastructure, ML/AI, and Data Science. Act as the primary point of contact for Talent Directors and senior tech leaders on new hire offers, internal equity reviews, leveling decisions, and out-of-cycle requests. Own compensation cycle execution f
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: MarTech Data Science Measurement empowers Airbnb to optimize marketing ROI by generating data-driven recommendations. We lead the way in defining and advancing best practices for measuring and optimizing marketing impact. We collaborate with Marketing, Finance, and Engineering to provide actionable recommendations and tools based on effective, timely, and granular measurements. Our team’s tenets are: Actionable: Deliver insights that drive confident business decisions. Impactful: Prioritize projects based on their expected value to Airbnb. Balanced: Adapt methods to business questions and data realities, acknowledging limitations. Rigorous: Maintain methodological integrity and quantify the sensitivity of findings. Innovative: Invest in advancing measurement science and developing new methods. Influential: Share learnings across Airbnb and the broader data science community. The Difference You Will Make: We are seeking an experienced (Contract) Sr. Data Scientist for a 24 month contract with deep expertise in marketing measurement, with a particular focus on Marketing Mix Modeling (MMM) and geo-based causal inference. The ideal candidate brings strong statistical intuition and hands-on modeling experience to quantify the incremental impact of Airbnb's marketing investments across channels and geographies. They are fluent in Python, comfortable working with Bayesian frameworks, and can translate complex measurement findings into clear, actionable recommendations for senior stakeholders. A Typical Day: Marketing Mix Modeling: Design, build, and maintain MMM models that est
NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,
From $259K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are seeking experienced machine learning engineers who thrive on solving complex challenges and designing scalable, ground breaking solutions. In this role, you will develop deep learning-based models for performance advertising business. Your work will lay the foundation to deliver effective performance ads to our users, and more business values to our advertisers. You will build innovative machine-learning solutions to power ad ranking algorithms, and personalized advertising experiences. With our ads system still in its early stages, this is a unique opportunity to shape and develop a world-class, ML-driven advertising platform from the ground up. You Will: Drive the design and implementation of machine learning solutions for ad ranking algorithms. Design and implement large scale recommendation models Author specs for new features and improvement Collaborate with other teams within Roblox to make sure we are building products with a community first approach. Balance researching new technologies with a practical approach to accomplish the research efforts into the Roblox products Communicate with the industry and commun
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Build the future of data. Join the Snowflake team. The Snowflake Machine Learning Platform team’s mission is to enable customers to bring their machine learning and deep learning workloads to Snowflake. Our customers want to build powerful models with the ever-increasing data in Snowflake but face several challenges including infrastructure optimizations, orchestration, performance, and security. The team aims to solve these challenges by building highly integrated platform solutions that are simple, secure, and enable end-to-end ML workflows. We are on an early journey to build the most scalable machine learning and data platform without sacrificing the benefits of a single platform and governance. We are looking for outstanding technical leaders who will join our ML Platform team to build the next-generation platform and play a pivotal role in this journey by understanding Snowflake’s core platform architecture and evolving it to enable state-of-the-art machine learning and LLM workloads. Join us to define strategies, set technical directions, design and execute, engage and deliver innovation, and unlock the power of AI for thousands of enterprise customers. This position is based in Menlo Park, CA, and Bellevue, WA. RESPONSIBILITIES : Help define and own the roadmap, wor
From $195.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Recommendation Systems are a key growth lever at Roblox, driving retention, engagement, and monetization for hundreds of millions of users. This role offers the unique opportunity to redefine how users search and discover everything from the most interesting immersive experiences and digital avatars in our Marketplace to personalized advertising. You will solve a diverse range of high-scale ranking, retrieval, and personalization problems across our platform. We combine cutting-edge research —including deep learning, generative AI, and reinforcement learning techniques— with large-scale engineering to bridge experimentation and production; you'll design algorithms that operate at massive scale and shape the next generation of recommender systems for user-generated content. Teams Hiring for This Role Search and Discovery: powers major recommendation surfaces—conducting cutting-edge research in generative modeling, multimodal and MLLM technologies, designing advanced agentic AI algorithms to solve business requirements while achieving technical breakthroughs. Safety, Alt Defense: architects a massive-scale detection engine that identifies recidivist bad actors across billions of account
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company”. We are looking to grow our company, and grow our teams with the smartest people in the world. What you’ll be doing: You will work with ground breaking technologies for the Tegra SoC and various NVIDIA embedded platforms Implement power and thermal management software features in Linux Kernel and user space Collaborate with power architects, hardware and software engineers on platform power estimation and optimization Optimize the software stack to improve performance, efficiency, and responsiveness for edge AI and robotics use cases. Focus on improving compute and memory utilization, reducing latency and power consumption, and tuning system-level performance to deliver reliable and scalable AI workloads across demanding real-world edge environments. What we need to see: MS in CS, CE, EE, Systems Engineering or related software/hardware engineering major, or equivalent experience 8+ years of software development experience with a significant focus on Linux Excellent C programming/debugging skills within Linux kernel and user space software Background with working on embedded systems and ARM processor specific System-level debugging experience and problem-solving skills Excellent communication skills Ways to stand out from the crowd: Understanding of the Linux power and thermal management features (schedule
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world. We are looking for an excellent Senior Engineering Manager to lead a large firmware engineering organization delivering end-to-end manageability firmware for NVIDIA's next generation Data Center Compute Systems. This role owns HGX product line and OpenBMC-based management firmware and MCU firmware components in data center platforms, including architecture, execution, quality, reliability, telemetry, and customer readiness. We are seeking an experienced senior leader with strong technical depth, broad system perspective, and a proven ability to lead large teams through complex product cycles. This role is onsite in Santa Clara, CA, USA. If you're creative and autonomous, we want to hear from you! What you'll be doing: Lead a large firmware engineering organization delivering OpenBMC based firmware and MCU firmware for next-generation Data Center Compute Systems. Own HGX platform as a lead for Firmware and System software readiness working across the organization. Define and drive the long-term firmware roadmap, balancing architectural innovation with product execution and delivery milestones. Drive architecture strategy across BMC, MCU, platform software, manageability, health management, and data center firmware interfaces. <spa
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li
Other cities to consider
More places hiring for this role
Get new senior deep learning framework communications engineer jobs in United States by email
Daily job updates · Unsubscribe anytime