NVIDIA is a global leader in high-speed computer vision, artificial intelligence (AI), and deep learning. Our team develops data engineering solutions that empower AI developers in autonomous vehicle (AV) domains to innovate quickly and effectively at scale. Are you ready to take on a senior technical role in building high-performance AI data pipelines? We seek an exceptional individual to design and optimize microservices and data pipelines to process massive volumes of AV data and enable seamless data mining and AI training. The ideal candidate will bring expertise in big data processing and distributed computing to create efficient solutions and overarching architectures for challenges such as video data curation, behavioral search, and AI dataset management. What you'll be doing: Scope and build tools, microservices, workflows, and distributed applications to accelerate data mining and AI training. Design and implement solutions for streaming, resilience, logging, security, authentication, workflow orchestration, and data management. Deploy AI models. Design and develop Retrieval-Augmented Generation (RAG) workflows enabling hybrid and agentic patterns. Analyze and operationalize complex distributed systems for speed-of-light performance. What we need to see: Experience developing high-performance, scalable software systems. MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field. Strong programming skills in Python or Golang Proficiency in key technologies like Kubernetes, Helm, Hive, Parquet, SQL, vector databases, e.g., Milvus. Strong architectural skills with a proactive, problem-solving mentality. Experience in data mi
Jobiba hiring network
Nvidia Manager Jobs
561 active opportunities Β· Updated for October 2026
Fresh results
15 shown
Explore current nvidia manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own. Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill, safety policy, or trace works the same no matter which harness is running. We want to enable agents that act on a person's behalf, governed and secure, continuously evaluated and self-improving. These agents coordinate and hand work off to each other, with identity and policy following every hop. They route and tune themselves across harnesses from live eval signals, and get better from their own production telemetry instead of waiting on a human to retrain them. Have you run agents on a harness like Claude or Codex and hit the walls that show up when they run for real, for days, against live systems β and wanted them to learn from it on their own? We're building the platform that solves those problems once, for every team. What you'll b
NVIDIA's Silicon Co-design Group (SCG) sits at a rare intersection: we own the full product development lifecycle, from early architecture definition through silicon bringup to product release. Our ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. If you want to see your work go from whiteboard to world-class silicon, this is where that happens. What You'll Be Doing: Architect and integrate system-level performance and power management features, controllers, and policies to optimize product efficiency across datacenter and client products . Build feature roadmaps to address low-power, low-noise, and performance-per-watt product needs through prototyping, use-case analysis, and cost/benefit trade-offs. Partner with architecture, ASIC, board/platform, software/firmware, and marketing teams to drive design decisions and debug complex issues. Track industry trends and market needs and translate them into forward-looking roadmaps that keep NVIDIA's products ahead of the curve. Lead debug efforts, develop workarounds, and support bringup , validation, manufacturing, and customer escalations. What We Need to See: <
NVIDIAβs DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data
NVIDIA is well positioned as the 'AI Computing Company', our GPUs being the brains that power modern Deep Learning software frameworks, accelerated analytics, modern data centers, and driving autonomous vehicles. We are looking for a Senior Software QA Test Development Engineer to join in the mission of crafting a distributed technology for all NVIDIA teams that remotely manage 10s of 1000s of resources in a simple and controlled fashion, allowing engineers to focus on engineering and automation, rather than being burdened by manual operational tasks. SWQA test developer engineers at NVIDIA are responsible for creating test plans, execution, and reporting, as well as developing scripts for test automation, designing and developing tools for the QA team, and developing integration tests for validation. As a test developer, you must identify weak spots and constantly design better and more creative test plans to break software and identify potential issues. You will have a huge impact on the quality of NVIDIA's products. The ideal candidate must have strong programming skills and hands-on experience using AI development tools to improve quality and productivity across the end-to-end QA workflow. This includes leveraging AI assistants for test automation, code generation, debugging, and enhancing testing efficiency. During the interview process, we will assess your ability to effectively use AI development tools and evaluate your programming capabilities to ensure you can deliver high-quality solutions. What youβll be doing: Architect, implement, and evolve scalable agentic end to end SWQA workflow, automated test frameworks, infrastructure, and tooling for complex software products. Define test strategy and quality gates across functional, integration, regression, reliability, and release-validation workflows. Build and maintain high-value automated coverage for Linux-based, co
NVIDIA is searching for a highly motivated, creative engineer with experience in system and AI software to join the GPU System Software team. As someone who is hardworking and passionate about their work, you will design key aspects of our production GPU kernel drivers, embedded SW and SOC platforms. You should demonstrate the ability to excel in an environment with complex software and hardware designs. What you'll be doing: Define, design, develop and verify features for our new SoCs platforms, collaborating with hardware engineers and fellow software engineers. Follow the new SOC platforms all the way through the development process to the customer products that are used throughout the world. Be heavily involved with the early firmware, performance, power management, and all of system software required to produce our world-class products. Have multiple opportunities to collaborate and communicate effectively with teams across the globe. What we need to see: BS, MS or PhD degree in Computer Engineering, Computer Science, or related degree, or equivalent experience. Strong C programming, C++, low-level driver, SOC system platform experience , and AI software design, arch and optimization. Familiarity with computer system architecture, microprocessor, and microcontroller fundamentals. Kernel experience with Linux, Android, Chrome, or Windows systems. Experience with complex
We are seeking a Front-End Integration Engineer for the NIC Silicon group. Join our team at NVIDIA's Networking business unit and be part of the innovative build and implementation of the next generation Network Adapter Silicon chips. Contribute to the development of powerful communication devices. What You'll Be Doing: Own and maintain Continuous Integration pipelines. Monitor, complete, and fix daily and nightly CI flows in front-end build areas. These include Lint, Synthesis, Equivalence Checking and Simulation. Automate EDA Flows: Build, develop, and maintain robust automation scripts using Python, Tcl scripting, and Shell to integrate EDA tools (Synopsys, Cadence) into automated build pipelines. Manage Perforce Integration: Maintain multi-IP branch/stream strategies, manage workspace specs, complete hardware build drops, and handle release labels/tags in Perforce (Helix Core). Support Engineering Teams: Act as the primary point of contact for RTL and verification engineers to debug flow crashes, bottlenecks, environment setups, and CI failure reports. Enforce Quality Gates: Implement automated check-in and change list triggers to run sanity checks, linting, and style enforcement before code is committed to main streams. Optimize Infrastructure: Monitor compute farm resources (LSF) and storage usage to reduce pipeline runtime and improve overall execution efficiency. What We Need to See: Experience: 2+ years of hands-on industry experience in ASIC/SoC front-end integration, CAD, or EDA flow automation. Education: Bachelorβs degree or Masterβs degree or equivalent experience in Electrical Engineering, Computer Engineering, Computer Science, or a related field. Perforce Expertise: Strong hands-on experience managing Perforce (P4 / Helix Core) streams, workspace specs, c
NVIDIA is seeking a creative and analytical Capacity Planner to join our Operations team and drive the scaling of our Boards & Systems manufacturing across all business units. In this cross-functional role, you will lead the weekly planning cycle, collaborate with Operations, Manufacturing, Supply Chain, Engineering, Finance, and Business Units, and leverage advanced analytical models to resolve capacity constraints, rationalize capital investments, and drive execution strategies for senior management. What You'll Be Doing Run the weekly capacity planning cycle for Boards & Systems, including forecast ingestion, supply planning alignment, Contract Manufacturer (CM) and testing reviews, and preparation of planning outputs. Consolidate and validate core planning inputs (18-month forecasts, NPI consumption, SKU mappings, CM commits, yield data) across enterprise applications like Anaplan. Analyze required versus available capacity, identify bottlenecks, and perform scenario analysis using data models to drive risk mitigation strategies. Lead capital investment, decisions based on capacity requirements, rationalize forecast changes, and optimize volume loading across CMs to meet target supply and revenue goals. Coordinate workflows across Supply Planning, Business Units, Contract Manufacturers, Production Planning, Purchasing, Manufacturing/Test Engineering, Finance, and NPI in a multi-program platform environment. Partner with IT teams to integrate planning models with enterprise systems, establishing standardized workflows, data pipelines, and toolsets. Present clear, data-driven execution plans, operational risk insights, and mitigation strategies to senior management. Align with cross-regional teams (US, Israel, APAC) to maintain seamless operational execution and standardized planning frameworks. What We Need to See 5+ years
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole β fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness. What you'll be doing: Drive rack-scale SW/FW architecture alignment across CSP engagements β including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration Drive technical work streams with CSP engineering teams on rack-scale system software β ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation Capture and synthesize CSP engineering feedback on rack-scale system software β health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior β champion that feedback into NVIDIA's architecture decisions Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development Identify cross-CSP patterns in rack-scale SW/FW iss
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services. What youβll be doing: Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale. Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation. Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data. Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems. Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation. What we need to see: BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services. Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges. <l
We are looking for a Solutions Architect to help customers and partners in South East Asia embrace NVIDIA technologies to build, deploy, and scale vision AI and video AI solutions. This is a highly technical role that requires deep expertise in VLM AI models and software engineering practices. We need a passionate, hard-working, expert and creative individual to help us pursue the many opportunities in this region. A Solution Architect is the first line of technical expertise between NVIDIA and our customers, as well as our partners. Your duties will vary from solutions design, training/workshops, troubleshooting, project coordination, industry and marketing speaking engagements, customer relationship management and more. What you'll be doing: Drive the adoption of NVIDIA's vision AI and video AI capabilities. Provide hands-on technical leadership to put VLM (Visual Language Models) and WFM (World Foundation Models) into applications for computer vision, document processing, and video analytics. Support solution development and technical validation around NVIDIA vision AI and video AI technologies Work with field, product, and engineering teams to translate customer requirements into deployable architectures, technical feedback, and roadmap input. Deliver demos, workshops, technical reviews, and partner enablement sessions that accelerate adoption of NVIDIA AI technologies. What we need to see: BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, Physics, Mathematics, Machine Learning, or a related field, or equivalent experience. 5+ years of proven experience in computer vision, vision AI, video AI, or world foundation models. Hands-on experience in building and scaling production VLM and multi-modal AI systems. Strong Python ski
NVIDIA is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as GB200. This role combines deep technical expertise in embedded firmware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment. What you will be doing: Design and develop firmware solutions for manageability and observability of data center servers. Actively participate in hardware bring-up activities, OOB firmware development, protocol stacks (Redfish, PLDM, MCTP, NSM) and hardware-software co-design for Cloud Service Provider deployments. Debug and troubleshoot NVIDIA GPU firmware issues, power management, performance, and thermal control problems for data center deployments, providing active support to CSPs. Partner directly with CSPs to deliver technical solutions, co-develop & co-debug features and optimizations, and provide support during new product introductions. Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments. Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation. What we need to see: Deep expertise in data center server architectures, HPC systems, and hardware-software co-design. Deep expertise in embedded firmware, server management controllers, and hardware bring-up with proven track record of shipping production BMC solutions Strong knowledge of DMTF protocols (Redfish, IPMI, PLDM, MCTP, SPDM), telemetry frameworks, and out-of-band management architectures Expert-level skills in C/C&
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Itβs a unique legacy of innovation thatβs fueled by great technologyβand amazing people. Today, weβre tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing whatβs never been done before takes vision, innovation, and the worldβs best talent. As an NVIDIAN, youβll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, we're not just transforming the world of computer graphics and AI; we're setting the stage for the future of autonomous driving. As a Lead Safety Architect, you will be at the forefront of our autonomous vehicle technology, ensuring its safety at scale. You will collaborate with the most innovative engineers and technologists to integrate safety measures into our latest DRIVE products. This role is paramount in achieving and exceeding NVIDIA's high safety standards, making your work both exciting and impactful! What youβll be doing: Representing NVIDIAβs functional safety strategy and architectures to the customer Working closely with customers to understand their functional safety requirements and system architectures and feeding those back into the development teams Assisting customers to safely integrate and validate our products in their systems and vehicles Supporting customer facing safety collateral Tailoring functional safety platforms and safety analyses for strategic customers You will be working closely with safety management, solution architects, sales and technical marketing teams to deliver state of the art products
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads β including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack β from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability β monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack β build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<
Get new nvidia manager jobs by email
Daily job updates Β· Unsubscribe anytime