Jobiba hiring network

Senior System Software Safety Engineer Jobs

7,292 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current senior system software safety engineer jobs. Use filters to narrow by work mode, employment type, experience and date posted.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by great technology—and outstanding people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. Our team is looking for Senior System Engineers specializing in self-driving vehicle technology. As a senior member you will be responsible for integrating and maintaining end-to-end software with Sensing/Perception/Localization/Planning/Control included, ensuring its stable and safe deployment in the field. What you'll be doing: Develop and maintain the application as well as vital tools for autonomous driving Diagnose and tackle real exciting problems to provide the world with the best driving experience Vertical stack performance optimization What we need to see: BS/MS or higher in computer science or a related engineering field Excellent C and C++ programming 5+ years of relevant proven experience Experienced in developing system software in user space with capability digging into kernel and even low-level hardware Good understanding of Operating Systems, threading, synchronization and parallel computing to build highly efficient applications Familiar with underlying parallel architectures, CPU/GPU/DLA/DSP Excellent analytical, communication and collaboration skills in international organization Ways to stand out from the crowd: Software development experience with CUDA Prior experience in: Autonomous vehicles, Robotics, Computer Vision and/or ML Understa

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxartificial intelligence
View job →
G
15 days ago

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

pythonlinuxai
View job →
HI
HP IQ
📍 San Francisco• Full-time• $179K – $252K/yr
15 days ago

Who We Are HP IQ is HP’s new AI innovation lab. Combining startup agility with HP’s global scale, we’re building intelligent technologies that redefine how the world works, creates, and collaborates. We’re assembling a diverse, world-class team—engineers, designers, researchers, and product minds—focused on creating an intelligent ecosystem across HP’s portfolio. Together, we’re developing intuitive, adaptive solutions that spark creativity, boost productivity, and make collaboration seamless. We create breakthrough solutions that make complex tasks feel effortless, teamwork more natural, and ideas more impactful—always with a human-centric mindset. By embedding AI advancements into every HP product and service, we’re expanding what’s possible for individuals, organisations, and the future of work. Join us as we reinvent work, so people everywhere can do their best work. About The Role HP IQ’s System Software team enables on-device experiences to take full advantage of our hardware capabilities.Our team collaborates with internal and external partners to develop unique solutions using cutting edge hardware, sensors, algorithms, and interaction models. If you enjoy solving complex, interdisciplinary problems with a world-class team, we'd love to hear from you! What You Might Do Lead, grow and mentor a high-performing embedded software team in a startup environment Serve as a technical leader and escalation point for complex hardware/software integration issues Communicate technical vision, risks, and progress clearly to cross-functional and executive stakeholders Own system-level strategy, architecture, and technical direction across multiple products Drive execution of low-level driver and framework software development in C and C++ across multiple devices and operating systems Oversee integration between hardware and software, ensuring alignment across cross-functional teams Guide performance, power, and thermal optimization efforts Balance hands-on technical invol

javaredislinux
View job →
N
12 days ago

We are hiring senior engineers to work on the CUDA driver, a core component of our platform for accelerating general purpose computation on the GPU. Our team delivers features and improvements to better realize the potential of NVIDIA hardware for a growing range of computational workloads, ranging from deep learning, scientific computation, and self-driving cars to video games and virtual reality! CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, switches and the operating system. What you'll be doing: As a member of our team, you will use your design abilities, coding expertise, and creativity to deliver the best Compute platform in the world. You will craft elegant solutions to exciting problems and craft the future direction of CUDA as you collaborate with your peers across NVIDIA. You will evangelize, architect, and implement new CUDA features You'll oversee and drive development efforts across multiple teams Collaborate with members of hardware architecture teams Help define forward-looking improvements to the CUDA APIs and programming model Design and maintain performance and precision modeling Write effective, maintainable, and well-tested code Develop code for multiple operating systems What we need to see: Bachelor of Science or Master of Science degree in Computer Science, Electrical Engineering, or related field (or equivalent experience) 15&#43; years of relevant systems software development experience Strong C programming skills </

artificial intelligenceai
View job →
E
11 days ago

Everpure (NYSE: P) has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in revenue, its first billion-dollar quarter, and accelerating growth into FY27. Our strategic agenda spans the companies defining the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data platform providers, and the broader AI ecosystem. This type of work—work that changes the world—is what the tech industry was founded on. So, if you're ready to seize the endless opportunities and leave your mark, come join us. THE ROLE We are seeking an experienced Platform Software Engineer for our Systems Software Team. You will be working as part of a dynamic team and will be responsible for designing, developing, and testing system software functionality for Pure’s upcoming platforms. The work spans the gamut of Systems software and you will have the opportunity to work in a wide range of areas and features ranging from Platform drivers to networking and storage layers. WHAT YOU'LL DO Plan and influence the lifecycle of new Hardware Platforms. Work on problems ranging from design, bring up, to deployment, upgrades and fleet level reliability. Participate in the full lifecycle of new hardware platforms from early bring up through manufacturing release. Work closely with peer teams to debug complex HW/FW of new server hardware, including CPUs, chipsets, and peripheral components. Debug complex HW/FW issues across x86, PCIe, NVMe, and networking using lab tools (oscilloscope, logic analyzer, JTAG) and kernel/driver traces. Design, implement and improve remote server management capabilities (e.g., using standards like Redfish) and enhance Reliability, Availability, and Serviceability (RAS) features. Design, write and maintain software components in C/C++, Python, Golang and RUST. Collaborate with vendors on requirements specification and follow through to system delivery. Work closely with hardware engineers, system architects, and o

pythonlinuxai
View job →
N
12 days ago

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable &#34;flight recorder&#34; for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as &#34;Hardware Fault,&#34; &#34;Software Bug,&#34; or &#34;Environment Issue.&#34; This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define &#34;signals of impending failure,&#34; enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6&#43; years in systems programming. Experience building automated

pythonkuberneteslinux
View job →
N
10 days ago

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA's Silicon Co-Design Group (SCG) leads the full product development lifecycle, from early architecture definition through silicon bringup to product release. The ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. This is where ideas become chips, and chips become products that define the state of the art — and we're building that future with some of the most motivated engineers in the industry. We're looking for a Senior Memory Systems Engineer to own HBM and LPDDR integration in sophisticated SoCs. This role covers the full stack, including silicon, package, embedded software, testing, and product development. The engineer will resolve the toughest system-level memory challenges throughout the process, building solutions that hold up at scale. What you'll be doing: HBM & LPDDR System Integration and Bringup: Drive HBM and LPDDR system integration, bringup, characterization, and debug for next-generation SoCs — taking memory subsystems through the full arc from first silicon to production-ready at scale. Full-Stack Memory Closure: <s

N
Nvidia
📍 Remote, Poland• Remote
11 days ago

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, passionate, and self-motivated, we want to hear from you! We are looking for an experienced networking software engineer. An awesome candidate is highly technical who is also comfortable with dealing with enterprise customers. You will join a team of Solution Engineers focused on the Mellanox Networking, DGX Platforms, Container Orchestrators, Deep Learning containers, and other Enterprise related system software. SW Solution Engineers spend approximately 50% of their time helping customers with their most complex problems and 50% of their time doing R&D related work. This individual should have proven grasp of datacenter and networking technologies, to provide comprehensive solutions for complex installations, maintenance, or operations for a broad scope of leading-edge networking products. What you'll be doing: Take ownership and drive customer issues with Ethernet or InfiniBand network adapter/DPU deployments from inception to resolution. Develop features and tools as part of solution engineering efforts to support all Enterprise Service offerings including but not limited to Networking products. Work with NVIDIA Enterprise customers and internal users to improve the availability, reliability, and overall experience of working with NVIDIA Networking products. Bring independent analysis, communication, and problem-solving to customer experience. Collaborate with engineering to document, recreate and solve issues. What we need to see: BSc in Computer Science, Electrical Engineering, Computer Engineering, or related field (or equivalent experience). 8&#43; years system software developm

REMOTEkuberneteslinux
View job →
N
1mo ago

NVIDIA is seeking a Senior Technical Program Manager to join the CSP Engagements team, focused on deep technical engagement with hyperscale cloud service providers for NVIDIA’s next‑generation datacenter systems such as Vera Rubin NVL72. This role is intended for experienced systems and embedded software leaders—including software engineering managers, technical leads, or senior architects—who have led datacenter server and platform software programs and can operate as a trusted technical partner to hyperscale CSP engineering teams. As a member of the CSP Engagements team, you will act as the primary technical engagement leader between NVIDIA’s system software organizations and CSP platform, system software, and AI teams, ensuring alignment, readiness, and successful large‑scale deployment of NVIDIA‑based datacenter solutions. What you will be doing: Lead deep technical engagements with hyperscale CSPs as the primary NVIDIA point of contact for system software, firmware, and platform readiness for NVIDIA datacenter products. Partner directly with CSP system software, firmware, and infrastructure engineering leaders to align on software architecture, bring‑up plans, deployment readiness, and production requirements for NVIDIA‑based server and rack‑scale platforms. Represent CSP technical priorities internally, advocating for customer requirements and tradeoffs across NVIDIA’s system software, firmware, hardware, silicon, and product teams are aligned to customer needs, timelines, and constraints. Own the end‑to‑end CSP engagement lifecycle, from early technical alignment and pre‑production readiness through large‑scale deployment, escalation management, and sustained production support. Drive bi‑directional technical communication: translating CSP system‑level requirements into actionable focus areas for NVIDIA engineering teams, while clearly communicating N

linuxartificial intelligenceai
View job →

We are seeking software engineers to work on next-generation graphics and computing products. Our charter is to build the most stressful set of applications a GPU or high performance computing server would see in its life cycle. The best candidates will have strong C&#43;&#43; programming skills, thorough knowledge of graphics concepts and algorithms, a solid foundation of systems software with emphasis on OS fundamentals, and a deep understanding of current generation PC/hardware architecture. Excellent communication skills and a dedication to meticulous engineering practices are a requirement. As a system software engineer, you will extensively use your knowledge of operating systems, algorithms, and computer architecture to provide robust and efficient solutions to validate and test next generation processors. What you'll be doing: Working closely with architecture, hardware and driver teams through the product development lifecycle of computing and graphics processors, as well as compute products. Responsible for crafting software tools and infrastructure required for new chip development, validation, and productization. You will assess new hardware features and architect manufacturing diagnostic tests using pre-beta CUDA and OpenGL extensions. This job will require an understanding of our hardware and software architectures. What we need to see: BS or MS degree in one of the areas of Electrical Engineering, Computer Engineering, Computer Science or equivalent experience 3&#43; years experience in a related hardware/software position Strong C/C&#43;&#43; programming skills Familiarity with P

🔔

Get new senior system software safety engineer jobs by email

Daily job updates · Unsubscribe anytime