Jobiba hiring network

Software Engineer Secrets Infrastructure Jobs

6,326 active opportunities ยท Updated for October 2026

Fresh results

15 shown

Explore current software engineer secrets infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.

N
1mo ago

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads โ€” including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack โ€” from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability โ€” monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack โ€” build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

pythonawsazure
View job โ†’
N
1mo ago

NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services. What you will be doing: Leading the overall architecture and design of our distributed storage service optimized for AI/ML Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes. Develop and maintain user-space applications, containers, Go-bindings, and CLI tools. Building features for a distributed storage service to enhance availability and reliability for large-scale deployments Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services. Automating distributed storage service end-to-end, including deployment, management, and monitoring What we need to see: Bachelorโ€™s of Science in Computer Science, or related field (or equivalent experience) with 8&#43; years of industry experience Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations Strong track record of delivering distributed services in a variety of distributed computing environments Experience in i

kubernetesartificial intelligenceai
View job โ†’

We are looking for a talented Senior System Software Engineer to help build NVIDIA Halos, our full-stack safety critical platform for physical AI: autonomous vehicles, humanoids, industrial robots, and intelligent machines that sense, decide, and act in the real world. This role is equally grounded in ISO 26262 functional safety engineering and embedded systems, rich operating systems, virtualization, accelerated computing. If you have a good understanding of System Software development on Real Time OS (RTOS), ARM architecture, Virtualization, and experience with ISO 26262 safety lifecycle, we want to hear from you! We are a leading artificial intelligence computing company and are paving the way with innovations in gaming, visualization, supercomputing, and autonomous vehicle and robotics safety platforms. NVIDIA gives automakers, tier-1 suppliers, automotive research institutions, and startups the power and flexibility to develop and deploy breakthrough artificial intelligence systems for autonomous vehicles. Our unified computing architecture makes it possible to train deep neural networks in the data center on the NVIDIA DGXโ„ข, and then seamlessly run them on NVIDIA's Halos platform inside the vehicle. Leading vehicle manufacturers, tier 1 suppliers, mapping and simulation companies, software and sensor providers, and startups around the world are developing on the NVIDIA Halosยฎ platform to deliver the best solutions for the new world of mobility. As a System Software Engineer, you will get an opportunity to work on high-performance NVIDIA edge AI platforms running complex mixed-criticality software stacks across QNX, Linux, hypervisors, safety partitions, accelerators, and real-time AI workloads, and work alongside industry experts in diverse teams and projects. <

linuxartificial intelligenceai
View job โ†’

By submitting your resume, youโ€™re expressing interest in our 2027 RDSS (Research and Development Substitute Services) program. Please confirm your eligibility with the local district office before applying the role. We are now looking for a Diagnostic NIC Software Engineer. The NPI Diag TW team is searching for Firmware Software Engineer to develop efficient secure and stable diagnostic firmware/software tools and scripts to enable factory build and debugging on NVIDIA products. In this role, engineers should be able to develop, maintain, debug CX/BF chip firmware/software and NBU Diagnostic software upon the next generation NVIDIA DGX/MGX servers, GPU baseboard, and PCIe cards. You will participate in a focused effort to develop and productize ground-breaking solutions that will be applied on many NVIDIA products. You'll find the work is exciting, fun, and meaningful. We have deadlines, customers, and competitions. We are the leading artificial intelligence computing company and are paving the way with innovations in gaming, visualization, supercomputing and self-driving cars. As a key member of our firmware team, you will be a key leader responsible for the Firmware of our DGX/GPU software stack. What you'll be doing: Be involved in the definition, architectural design, and development of security Diag tool and test coverage planning for NVIDIA DGX products with an opportunity to craft its future. Assist with defining and making sure the software development process meets security and performance standards. Design and/or make recommendations for Network solutions that apply to our software to satisfy DGX/GPU server guidelines and requirements. Design Diag tool and scripts to make sure DGX/GPU products go for production smoothly. Design Diagnostic software to help validate GPU/CPU/NIC/DPU HW with good performance and quality. <

pythonlinuxartificial intelligence
View job โ†’
N
1mo ago

We are looking for a highly motivated AI/ML Software Engineer to join the Enterprise Agentic AI Platform team within IT. You will work closely with Business Analysts, and Engineering teams to design, develop, and deploy enterprise AI solutions that improve productivity and automate business workflows across Engineering, Operations, and Manufacturing. What you'll be doing: Design, develop, and deploy Agentic AI applications using Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI orchestration frameworks. Build scalable AI services and reusable components integrated with enterprise applications such as PLM, SAP, and other business systems. Collaborate with business and IT teams to translate business requirements into AI-driven solutions. Develop secure, scalable APIs and enterprise integrations to enable intelligent workflows and automation. Improve AI solution quality, performance, and reliability through prompt engineering, evaluation, and continuous optimization. Partner with cross-functional teams throughout the Software Development Lifecycle (SDLC), from solution design through deployment and production support. What we need to see: Bachelor's or Master's degree in Computer Science, Information Technology, AI/ML, or a related field. 6&#43; years of software engineering experience with strong proficiency in Python and backend application development. Hands-on experience with Generative AI, LLMs, RAG, AI agents, REST APIs, and cloud-native application development. Experience integrating enterprise applications and building scalable, production-ready software solutions. Strong analytical, problem-solving, communicatio

pythonazureai
View job โ†’

Lead Java Software Engineer - Developer Company: The Boeing Company The Boeing Company is looking for a Lead Software Engineer - Developer to join the Advanced Ground Architecture team located in Herndon, Virginia, Colorado Springs, Colorado, Mesa, Arizona, Seal Beach, California and El Segundo, California. This position will focus on supporting the Boeing Defense, Space & Security (BDS) Software Engineering organization. The Advanced Ground Architecture (AGA) software team is dynamic group of software engineers creating the future of Ground support with the extensibility and adaptability to be used across ALL Boeing programs. The software team is executing this vision through modern software technologies (Java, ReactJS, python, CI/CD pipelines) and methodologies (Scaled Agile). AGA is looking for self-motivated high performers to lead and execute the large scope of Java development needed for the program's vision. The ideal candidates will provide Java software development leadership for the Advanced Ground Architecture program, specifically the System Planning and Scheduling team. These positions will be responsible for accelerating complex java development, shepherding peer reviews for code merge requests at a faster cycle time, providing sub system leadership on key components, mentoring early and mid career software developers in AGA to be higher performers, and assisting P6 and Chief Engineering on architecture and software system engineering duties and decisions. These responsibilities are in addition to monitoring overall cost and schedule performance to meet or exceed expectations. Position Responsibilities: Oversees the design, deve

pythonjavareact
View job โ†’

Experienced Simulation Software Engineer - Training Systems Company: The Boeing Company The Boeing Company is looking for an Experienced Simulation Software Engineer - Training Systems to join our Government Vehicle Health Management Systems (GVHMS) team in Hazelwood, MO . The Government Training Team develops innovative simulation software solutions that advance pilot readiness and mission success. This Software Engineer role involves designing, architecting, and developing simulation models, virtual environments, and frameworks, while collaborating with stakeholders to optimize overall simulation performance. Responsibilities include simulation validation, integration, and modernization of legacy software within a secure Agile development environment. The role requires strong expertise in C&#43;&#43;, with additional knowledge of Python, containerization, and cloud-based technologies. Familiarity with emerging software engineering methods and prior aviation or engineering experience are highly valued for contributing to this fast-paced, mission-focused team. At The Boeing Company, we innovate and collaborate to make the world a better place. From the seabed to outer space, you can contribute to work that matters. Weโ€™re committed to fostering an environment for every teammate thatโ€™s welcoming, respectful and inclusive, with great opportunity for professional growth. Find your future with us. Position Responsibilities: Designs, architects, and develops simulation models, simulation visualizations, virtual environments/platforms, and frameworks to enhance test performance, safety, and durability of software and hardware throughout the entire product lifecycle Partners with stakeholders to identify simulation r

pythonawsazure
View job โ†’
H
Hp
๐Ÿ“ Texasโ€ข $147.1K โ€“ $230.9K/yr
1mo ago

Principal Embedded Firmware and Software Engineer Description - We are seeking a Principal Embedded Firmware & Software Engineer to lead the design, development, and debugging of embedded software and firmware for computer systems. In this role, you will combine deep, hands-on engineering expertise with system-level technical leadership to ensure seamless integration between software and hardware components, delivering reliable and efficient system performance. You will collaborate closely with cross-functional teams including hardware engineers, software developers, QA, and product managers to bring high-quality products to market. Responsibilities Provide technical leadership for the architecture, development, security, integration, debugging, validation, and deployment of embedded firmware and software including BIOS/UEFI, EFI applications and drivers, embedded controllers, and RTOS-based systems. Analyze hardware and system architectures to define firmware requirements, dependencies, interfaces, integration strategies, and validation approaches. Troubleshoot and resolve firmware issues by designing and implementing enhancements, updates, and programming changes across firmware subsystems. Define and drive firmware integration, verification, and validation strategies, including automated testing, regression testing, and continuous integration . Develop and improve engineering tools and automation using Python and other appropriate technologies for development, debugging, testing, analysis, and validation. Advance CI/CD and DevSecOps practices for embedded development to improve engineering velocity, quality, traceability, and release confidence. Evaluate and apply AI-assisted software development and engineering tools where they can improve developer productivity,

pythongitlinux
View job โ†’

NVIDIA has transformed computer graphics, PC gaming, and accelerated computing for more than 25 years through exceptional technology and the people who build it. In semiconductor manufacturing, our role is to enable the ecosystem, not compete within it. We partner with fabs, equipment manufacturers, and software providers to make inspection, metrology, and manufacturing intelligence dramatically faster on the NVIDIA platform. Our team builds the software that makes this possible: models, adaptation and evaluation workflows, and deployable inference capabilities that partners integrate into their own tools. We work in environments where labeled data is limited and proprietary, distributions shift across tools and fabs, production budgets are tight, and software must operate inside air-gapped facilities. Weโ€™re seeking a Principal Systems Software Engineer for Semiconductor Inspection in Santa Clara. This is a hands-on architect role: you will define the approach, build it, evaluate it, and demonstrate the results. You will work across computer vision, time-series modeling, multimodal AI, anomaly detection, model adaptation, evaluation, and production inference. Success means technology that a fab or equipment vendor can integrate, operate, and trustโ€”not only a successful internal demonstration. What youโ€™ll be doing: Define and prototype AI system architectures spanning optical and e-beam inspection, wafer and mask inspection, metrology, defect review, equipment signals, and process data. Advance world foundation model capabilities for semiconductor manufacturing, including vision, time-series and multimodal representation learning, model adaptation, domain transfer, and data-scarce defect understanding. Develop workflows for defect detection, classification, localization, segmentation, nuisance filtering, ADC, AD

pythonmachine learningai
View job โ†’

NVIDIA's Deep Learning GPUs have ignited modern AI โ€” the next era of computing โ€” with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as โ€œthe AI computing companyโ€. We are growing our company and the team with the smartest people in the world. We are looking for extraordinary Software Engineers to develop and productize NVIDIA's DRIVE OS software. As a member of NVIDIA's Solution Engineering team, you will adapt DRIVE OS solutions to various car platforms equipped with different sensors. We are looking to hire Senior System Software Engineer โ€“ AUTOSAR. Ideal candidate will have very strong programming skills, a good grasp of HW & SW Architectures, a solid exposure to AUTOSAR & related architecture, tools and frameworks. What you will be doing: Participate and provide inputs and recommendation into AUTOSAR Architecture evolution with design choices, tools and methodology Architectural explorations on both SW and HW fronts which include feasibility studies, quick prototyping, profiling, safety studies, data analysis and presentation of results Influence next-gen HW architectures and SW Architecture and design Drive complex technical issues to closure that may occur interacting with cross-teams What we need to see: BS/MS, or equivalent experience 5&#43; years of experience Strong programming skills in C/C&#43;&#43; and scripting skills in Perl, Python etc Good experience and com

Job Details: Job Description: In this role, you'll build software capabilities to automate the build, test, and deployment of Intel's Process Design Kit (PDK). A PDK is a collection of artifacts representing Intel's semiconductor process, used by product designers to model, implement, and verify Intel's mobile, desktop, and server products before manufacturing. You'll be a member of the Design Technology Platform organization, working closely with teams in the United States, Bangalore, and Penang to build world-class DevOps infrastructure that shapes the future of deploying PDKs to silicon product development teams at Intel. Qualifications: Bachelor's or Master's degree in Computer Science, Computer Engineering, or another closely related field. 7-12 years of experience with a Bachelor's degree or Master's degree in software development and engineering. Excellent Python programming skills. Expert-level knowledge of pytest concepts. Expert-level experience with GitHub-based Jenkins CI/CD deployment, enablement, and debugging. Proven knowledge of Agile software engineering practices, IT environments, and DevOps principles and processes. Excellent debugging and problem-solving skills. Excellent written and verbal communication skills; ability to present complex issues with clarity to drive decisions. Must have hands-on experience with AI tools and demonstrated experience building AI agents/tools. Exposure to VLSI PDK domain is preferred. Strong team player with proven ability to collaborate effectively across cross-functional and geographically distributed teams. Job Type: Experienced Hire Shift: Shift 1

pythonaijenkins
View job โ†’
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access. Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows. Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems. Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification. What we need to see: BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5&#43; years of software engineering experience. Strong foundation in algorithms, data structures, distributed systems, and practi

pythonjavaaws
View job โ†’

We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based application crafted to provide our storage customers the capabilities to handle and supervise our distributed storage infrastructure. Our team is continually dedicated to acquiring and implementing ground breaking technologies to overcome obstacles and innovate solutions for improving our ability to handle large clusters of machines efficiently. What You Will Be Doing: Maintain and develop Kubernetes operators and our Container Storage Interface (CSI) plugin. Develop a web-based solution that manages, operates and monitors our distributed storage. Work closely with other teams to define and implement new APIs. What We Need to See: B.Sc., M.Sc. or Ph.D. in Computer Science, or related discipline, or equivalent experience. 8&#43; years of experience in web development ( both client and server ) Proven experience with Kubernetes (K8s), including developing or maintaining operators and/or CSI plugins. Experience scripting with Python, Bash or similar. Experience with nodejs is a must At least 5 years of experience working in a Linux OS environment Youโ€™re smart and a quick learner You do what it takes to get the job done Passionate about coding and big challenges Ways to stand out from the crowd: NodeJS for the server side: dominant modules are async & express . Kafka, MongoDB, K8s JavaScript frameworks: React, jQuery, c3j

javascriptpythonreact
View job โ†’
N
1mo ago

The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys

pythonjavakubernetes
View job โ†’
N
1mo ago

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters. Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available Triage and root-cause performance issues reported by our customers Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information <li

pythondockerkubernetes
View job โ†’
๐Ÿ””

Get new software engineer secrets infrastructure jobs by email

Daily job updates ยท Unsubscribe anytime