Jobs in United States

Linux System Administrator in United States

217 active opportunities · Updated October 2026

Explore current linux system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

SCG sits at the crossroads of design, architecture, marketing, and productization—owning the journey from the architecture stage through final product definition across Gaming, Datacenter, Automotive, and Embedded markets. As a System Verification CoDesign Engineer, you will work on system-level speed features, develop the verification collaterals and automation infrastructure to characterize and validate them, and lead debug of the complex silicon issues that stand between a program and on-time shipment. This is a hands-on role for an engineer who combines deep technical craft with the drive to compress cycle time using modern tooling—including AI—without losing rigor. What You’ll Be Doing: Collaborate cross-functionally with system architects, hardware, firmware/software, process/reliability, and operations teams to co-design system-level speed features and deliver industry-defining products. Understand system level behavior and speed reliability margins, bounding box constraints and identify solutions that optimize margins . Translate hardware features and architectural requirements into verification techniques that achieve full coverage across testing flows. Perform closed loop validation by correlat ing silicon behavior against timing simulation and design expectations; provide actionable feedback to improve future designs. Define, prototype, and refine pre- and post-silicon bring-up flows to ensure

PythonLinuxAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads — including hyperscale multi-node training, inference, cloud gaming, and cloud functions. This role spans the full lifecycle of our SDN stack — from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response. What you'll be doing: Design and develop next-generation multi-tenant cloud SDN control and data plane software (OVS, OVN, OpenFlow) Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes Drive upstream contributions to OVN-Kubernetes and related open-source projects Develop software for network observability — monitoring, telemetry, intelligent metering, and performance analysis Operate and support OVS-OVN based SDN solutions in large-scale NVIDIA AI Cloud environments Own end-to-end observability for the SDN stack — build and maintain monitoring, alerting, distributed tracing, and dashboarding to ensure real-time insight into network health, performance, and tenant SLAs Design, enhance, and maintain CI/CD pipelines (GitLab) across Linux host networking, OVS, OVN, and Kubernetes CNIs Implement GitOps approaches or related experience for secure, seamless integration with cloud infrastructure Drive reliability through incident management, resource monitoring, and performance tuning<

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI's mission is to ensure that artificial general intelligence benefits all of humanity. The Consumer Devices team is building a new generation of AI-powered products that seamlessly integrate hardware and software to create intuitive, transformative experiences. We bring together experts across embedded systems, machine learning, hardware, design, and product engineering to develop products at the intersection of AI and consumer technology. About the Role OpenAI is seeking a System Performance Engineer to profile, benchmark, and optimize performance across our embedded hardware products. In this role, you will work across operating systems, applications, camera and vision, graphics, and platform teams to define product KPIs, build performance tooling, and drive optimizations from early lab characterization through product launch and real-world usage. You will help establish the performance standards that shape the user experience of our products, ensuring they remain responsive, efficient, and reliable throughout their lifecycle. This role requires deep expertise in embedded or high-performance systems, strong operating systems fundamentals, and hands-on experience debugging under tight latency, power, and memory constraints. This role is based in San Francisco, CA. We use a hybrid work model of four days per week in the office and one day working remotely. Relocation assistance is available for new hires. In this role, you will: Develop system performance benchmarks, methodologies, and policies to evaluate end-to-end product behavior. Profile and analyze performance across key product use cases and workloads using custom and industry-standard profiling tools. Partner closely with engineering teams to identify bottlenecks and drive performance optimizations across the software stack. Define high-level product KPIs and establish measurement frameworks to measure launch readiness and monitor performance throughout the product lifecycle. Measure, re

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. Role summary We are seeking a Networking Operating System Firmware Engineer to help bootstrap and scale the switching layer of our AI supercomputers. In this role, you will build and maintain custom NOS images from scratch, using open source components from SONiC, SAI, FRR, and related networking stacks while working across the Linux kernel, switch ASIC SAI/SDKs, platform drivers, control-plane services, and orchestration layers. This is a software engineering role that requires a deep understanding of networking, NOS internals, switch hardware, and production systems. You will design, implement, test, and debug production NOS software across platform drivers, routing and control-plane state, ASIC programming, observability, and fleet integration. The engineer in this role should be able to work through ambiguous, open-ended technical problems and drive feature development across software, hardware, and vendor boundaries. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Design, develop, and maintain custom NOS images for large-scale AI fabrics, using open source components from SONiC, FRR, and related networking stacks. Integrate, build and configure Linux kernel components, device drivers, switch ASIC SDKs, and SAI layers. Bring up new switch platforms, including thermal and fan control, power monitoring, transceiver management, watchdogs, OSFP CMIS, L

PythonAWSCI/CDLinux
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are looking for a 100% hands-on Storage Services Software engineer to join the block storage group. You will be a member of a team that builds the next-generation block storage capabilities and architects a proprietary distributed file system solution from its inception. You will work closely with a variety of teams and architects including the networking team, and external customers. You will take part in defining the software architecture and implementation of the most advanced storage services! Services that will need to meet extreme performance and scalability demands! We have crafted a team of extraordinary people stretching around the globe, whose mission is to push the frontiers of what is possible today and define the platform of tomorrow. At NVIDIA, we work, think and learn as a team. We thrive in a deeply strong environment, and we're passionate about a culture that demands innovation and the highest standards. The rewards are sweet and include collaborating with some of the smartest people in the industry, an aggressive compensation plan that rewards top performers, and the opportunity to work on products that transform the way people work and play. What you’ll be doing: 100% hands-on coding role in C language, Kernel and Userspace Access advanced AI tools and a token budget for code development provided by NVIDIA, the world's AI factory leader. Research, design, implement and test, new and existing, distributed storage services and features of NVIDIA’s block and file storage solution, in both Host and DPU environments. Acquire understanding of the algorithms, the technicalities and the interaction with other components across NVIDIA’s block and file storage ecosystem. Analyze and solve challenging bugs and customer cases in la

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. Our team works across custom silicon, embedded systems, operating systems, and cloud services to build reliable consumer devices and the platforms behind them. We connect kernel development with the broader software stack to deliver complete product capabilities. About the Role As an Operating Systems Engineer focused on the Linux kernel, you will design, develop, and maintain the kernel capabilities that underpin OpenAI’s consumer devices. You’ll bring deep expertise in one or more Linux kernel subsystems and carry solutions through the higher-level software stack. Your ownership will extend into the userspace services, libraries, tools, and interfaces needed to deliver complete product features. You’ll shape the boundaries between kernel and userspace, make design decisions across the stack, and see your work through development, integration, and production. In this role, you will: Build kernel capabilities: Design, implement, and maintain Linux kernel subsystem changes that support device capabilities and product requirements. Own features across the stack: Choose appropriate kernel and userspace boundaries, and build the interfaces and supporting components needed to deliver reliable features in shipped products. Debug complex system behavior: Use tracing, profiling, instrumentation, and diagnostic tools to resolve correctness, concurrency, p

LinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

We are looking for a disciplined and dynamic, Lead System Engineer – compute blade and rack Validation to join our growing compute rack validation team. As a diligent leader in Systems Engineering, you will drive multiple aspects of validation throughout the life cycle of the program. In this high visibility position, you will be part of a leading team to innovate and improve system bring-up and enablement abilities, as well as silicon and system validation to deliver the highest quality, industry leading technologies to market. Your technical leadership skills, validation and debug expertise will be necessary towards product development, definition, root cause and resolution. Your agility and collaborative approach will be essential to work within System Validation & other engineering teams (System Architects, SoC and Rack FW etc). The technical leader will be driving keys areas of system validation including leading first silicon & system bring-up (nodes and rack level systems) - rack level systems and blades will be based of ARM server architecture. Candidate will be immersed in challenging system enablement work, system validation (end-to-end) methodology, tests development and execution as well as triage/debug of critical issues to meet critical program milestones at POR quality. The candidate will also be a key contributor to state-of-the-art HW and lab capabilities for Grapchore’s system engineering. The candidate should be able to work in a global environment while maintaining a synergetic culture. Primary Responsibilities: Lead the systemenablement (including first silicon and other FW components) to ensure system capabilities are brought up as per plan of record and system architecture spec. Drive organization wide methodology for Firmware integration and best known configuration (HW/FW/SW) usage model by leading the release of deployment ready solutions. Develop key methodologies, lab HW and system SW capabilit

LinuxAIGoExcel
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver communication libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. DL and HPC applications have a huge compute demand already and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! We are looking for a technical leader to manage our NVSHMEM and UCX libraries. This is an outstanding opportunity to push the limits on the state-of-the-art and deliver platforms the world has never seen before. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision? What you will be doing: Lead, mentor, and grow your library engineering team and be responsible for the planning and execution of projects as well as the quality, and performance of your libraries. This is a technical leadership role so you will participate in feature design and implementation. Interact with internal and external partners and researchers to understand their use cases and requirements. Collaborate with engineering teams, program and product management, and partners to define the product roadmap. Continuously review and identify improvement opportunities in established processes, infrastructure, and practices to ensure the teams are executing in the most efficient and transparent manner. What we need to see: 10&#43; overall years of experience in the software industry with specialization in HPC networking or system software. 4&#43; years of management experience. BS, MS, or Ph.D. in C

🔔

Get new linux system administrator jobs in United States by email

Daily job updates · Unsubscribe anytime