About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role We're looking for an Engineering Manager to lead a team of highly experienced engineers building the infrastructure that powers Modal's serverless GPU platform. This is a hands-on leadership role — expect to split your time between technical contribution and people management depending on what the team needs. You'll set direction, remove blockers, and build a strong engineering culture as your team tackles hard problems in distributed computing, large-scale data handling, and performance optimization. Who You Are You're an experienced engineering leader who stays close to the work and builds alongside your team when it counts. You earn trust through technical depth, not title. You communicate clearly, help strong engineers move fast without cutting corners, and stay calm and pragmatic under pressure. You care as much about how your team gets to an answer as the answ
Jobs in United States
Linux System Administrator in United States
217 active opportunities · Updated October 2026
Showing
15 jobs
Explore current linux system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally. What You Will Be Doing: Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can b
Senior Software Engineer Mission Systems Software Developer Company: The Boeing Company The Boeing Company has an exciting opportunity for a Senior Software Engineer Mission Systems Software Developer (Level 5) to join Boeing Defense, Space & Security (BDS) in Daytona Beach, Florida . BDS is a global leader in the development, production, maintenance and enhancement of fixed-wing and rotary wing aircraft, commercial and government satellites, human spaceflight programs and weapons. Key markets include aeronautics, space and weapons. Core capabilities are in development, production and mission enabling upgrades of integrated solutions. BDS delivers the most digitally advanced, simply and efficiently produced and intelligently supported solutions to its customers. Daytona Beach is the company's newest state-of-the-art facility focused on engineering excellence. Florida offers no state income tax and a variety of other desirable personal and financial benefits. This position will include developing software and software tests throughout all phases of the software development life cycle (requirements, architecture, implementation, and verification). The software engineer will develop software in a Continuous Integration / Continuous Deployment (CI/CD) DevSecOps software build pipeline using an agile methodology focused on code quality, security and automation. Position Responsibilities: Develop Participates in use case development of software requirements Ability to lead activities to develop, document and maintain software architecture, requirements, algorithms, interfaces and designs for mission systems software Leads development of code and integration of complex Mission Systems softwar
The Undersea Systems Division at Leidos currently has an opening for a Technical Deputy Program Manager to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Technical Deputy Program Manager and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, a
The Undersea Systems Division at Leidos currently has an opening for a Technical Program Lead to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Technical Program Led and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, action-item logs, and
Job Details: Job Description: The Role and Impact Join Intel's Design Technology Platform (DTP) organization and become an integral part of a highly skilled team driving the development of cutting-edge technologies. As an EDA Tools Software Engineer, you will design, develop, test, and debug software tools and methodologies that directly support Intel's hardware product design and manufacturing processes. Your work will enable design teams to innovate faster, meet complex manufacturability constraints, and deliver industry-leading products to market. By contributing to Intel's advanced process technologies, you will play a pivotal role in achieving leadership in design automation and fostering a culture of innovation. Key Responsibilities Design, develop, and debug Fill components of Process Design Kits (PDKs), Provide algorithmic solutions using tools such as Calibre, ICV, and Pegasus to address density deficiencies and ensure manufacturability compliance. Collaborate with process developers, design rule owners, and end users to define requirements and implement state-of-the-art solutions. Automate workflows to enhance efficiency, accuracy, and deployment across design teams. Develop comprehensive test cases and systems to validate software tools and ensure seamless integration with various design methodologies. Innovate and drive advancements in EDA tools and methodologies to meet the evolving needs of Intel's technology ecosystem. Partner with cross-functional teams and external EDA vendors to define and implement new tool features and methodologies. Maintain documentation, including training materials and user guides, to support customers and ensure robust deployment.
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: We are looking for a strong technical lead to guide the engineers designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. You'll lead the team responsible for Modal's machines layer: the fleet of bare metal and cloud hosts that every Function, Sandbox, and training job runs on, and the control plane that provisions, images, monitors, and repairs them. You'll own the full lifecycle of a machine, from accepting and benchmarking new hardware from a growing set of providers, to network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation of unhealthy hosts. You'll manage a team of 3–8 engineers while staying hands-on across the stack which involves BMCs, firmware, PXE, bootloaders, Linux networking, drivers, and distributed control-plane services, and you'll shape our long-
NVIDIA's GeForce Now, the next-generation gaming service powered by NVIDIA GPUs in the cloud, transforms a Mac, any PC, or just a mobile device into a high-performance gaming rig. GeForce NOW automatically keeps games up-to-date, and users around the globe can instantly stream the latest games in high-definition resolution at the lowest latency for the smoothest gameplay. Just click and play! Visit us at https://www.nvidia.com/en-us/geforce-now. In addition to gaming, the state-of-the-art low-latency streaming technology has expanded to a new range of applications, including augmented and virtual reality, artificial intelligence, and robotics. We are now looking for a Senior Software Engineer – Streaming with strong C++ skills and a deep interest in streaming technologies to join a team of highly skilled and motivated software engineers who help build the next generation of applications with media and data streaming capabilities. Now, are you passionate about driving this technology to its edge? Do you understand various streaming protocols? Can you solve complicated problems and propose innovative solutions? Then, we are keen to hear from you. What you’ll be doing: Design, develop, optimize, and debug C++ software for ultra low-latency streaming systems. Build and improve media streaming pipelines using technologies such as WebRTC, GStreamer, RTP/RTCP, and related protocols. Analyze CPU, memory, networking, media pipeline, and scheduling behavior across real-world workloads. Work across Windows, Linux, QNX, applications frameworks, and embedded platforms. Define and implement KPIs for networking quality, streaming quality, and user experience. Integrate streaming softwa
NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE Container runtimes were designed for general-purpose software workloads. AI inference is not a general-purpose workload. Running large models at production scale exposes cracks in every layer of the container stack: runtimes unaware of GPU memory constraints, images that take minutes to pull when a model needs to scale to thousands of replicas, and isolation mechanisms that weren't designed for the multi-tenant serving environments that production AI requires. The tools the industry has relied on for a decade weren't built for this, and patching around those limitations at higher layers only goes so far. Baseten owns the entire pipeline, from the moment a developer pushes a model to the moment a request gets a response. That vertical ownership means we can fix these problems at the root. The Runtime Fabrics team is doing exactly that: purpose-building the container runtime and storage layers for AI inference workloads, led by some of the world's top containerd maintainers. As Engineering Manager of the Runtime Fabrics team, you will lead this work, setting technical direction, growing a world-class team of systems engineers, and ensuring the team's output shapes not just Baseten's infrastructure but the open-source container ecosystem at large. If you've contributed to containerd, runc, or related OCI projects and are ready to lead a team solving some of the hardest problems in infrastructure today, we'd love
About the Team The GPT Infrastructure team builds systems that turn advances in model inference and optimization into reliable production capabilities. We enable OpenAI workloads to be qualified and optimized across new accelerator platforms without requiring a one-off port and tuning effort for every hardware target. Our work spans distributed systems, model execution, compilers and runtimes, performance engineering, secure partner integrations, evaluation systems, and developer tooling. We build the infrastructure that makes optimization workflows automated, reproducible, and trustworthy. About the Role We are seeking a software engineer to help build the platform that qualifies and optimizes inference workloads across heterogeneous compute environments. You will develop both OpenAI-hosted services and secure partner-side software for running long-lived optimization workflows. These workflows generate candidate kernels, runtime configurations, and serving-stack changes; compile and execute them on target hardware; verify their correctness; measure their performance; and use the results to guide further optimization. You will work across model architecture, distributed execution, compilers, runtimes, networking, and accelerator systems. A central part of the role is turning research prototypes and one-off hardware bring-up efforts into reliable, reusable infrastructure with clear contracts, reproducible results, strong observability, and well-defined security boundaries. Key Responsibilities Design, build, and operate APIs and control-plane services for long-running workload qualification and optimization campaigns, including scheduling, retries, checkpointing, resource budgets, and observability. Build secure partner-side execution and evaluation software that can compile, run, verify, profile, and benchmark candidate artifacts on accelerator hardware. Integrate model workloads, hardware profiles, compiler toolchains, runtimes, serving engines, and distributed-exe
About the Team The Consumer Products team at OpenAI builds end-to-end hardware and software systems that bring AI into the physical world. We work at the intersection of custom silicon, embedded systems, operating systems, and cloud services to deliver reliable, production-ready devices at scale. Within Consumer Products, the camera stack is a critical sensing component. The team partners closely with electrical engineering, silicon vendors, systems, and higher-level perception and product teams to bring up new hardware, stabilize capture pipelines, and ensure camera systems are robust, debuggable, and ready for real-world deployment. This work spans early prototypes through production, with a strong emphasis on correctness, repeatability, and long-term reliability. About the Role As a Camera Firmware Engineer, you will own low-level camera enablement on custom hardware—from early board bring-up through stable production capture. You will develop and maintain the firmware and software that makes camera sensors reliable, controllable, and debuggable, forming the foundation for higher-level camera pipelines and product features. This role is highly hands-on and systems-oriented. You will work close to the hardware, diagnose real-world timing and integration issues, and build tooling that accelerates iteration across the entire camera stack. This role is based in San Francisco, CA. We follow a hybrid work model with four days per week in the office and offer relocation assistance to new employees. In This Role, You Will Bring up new camera sensors and modules on prototype and production boards, including link stability, sensor control, and correct power, reset, and clock sequencing. Develop and maintain low-level camera software, including sensor drivers, board configuration, and camera subsystem integration across hardware revisions. Enable and validate core capture paths for development and production, including RAW capture for debugging, still capture, and hardware-
Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys
Other cities to consider
More places hiring for this role
Get new linux system administrator jobs in United States by email
Daily job updates · Unsubscribe anytime