Jobs in United States

Linux System Administrator in United States

217 active opportunities · Updated October 2026

Explore current linux system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $156K/yr

Quick readStrong listing-quality and freshness signals

The Team We are Datadog’s in-house product experts. The Technical Solutions team enables Datadog’s worldwide growth by educating potential partners and ensuring that our integration ecosystem is high-performing, secure, and valuable. Partner Technology Solutions Engineers (TSEs) are the technical bridge between Datadog and our third-party developer community. We act as consultants, helping partners build world-class monitoring solutions on the Integration Developer Platform (IDP) . The Opportunity Datadog is looking for a Partner Technology Solutions Engineer to join our fast-paced team. You will be the primary technical contact for our partners, guiding them through the entire integration lifecycle—from initial architectural design to final publication on the Datadog Marketplace. This is a unique role that combines deep technical troubleshooting with high-level consulting and platform advocacy. You will work directly with external developers and see your contributions immediately reflected in the Datadog ecosystem. You Will Act as the technical lead for partners, advising on OAuth flows, log pipelines, OpenTelemetry, and agent-based vs. API-based configurations Perform architectural assessments and deep-dive code reviews for partner integrations in the integrations-extras and marketplace repositories, ensuring they meet our Quality Rubric Solve complex technical challenges for partners via Zendesk, Slack, and dedicated technical consultations Identify friction points in our Integration Developer Platform (IDP) and partner with our internal Product and Engineering teams to build a better developer experience Maintain public-facing developer documentation and internal tracking systems ( JIRA ) to ensure transparency and scale You Are A technical expert with 3+ years of experience in a technical role (Support Engineering, Solutions Architecture, or Software Development) Proficient in at least one language (Python or Go preferred) An observability enthusiast who unders

PythonLinuxAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads. The Hardware Operations team is responsible for ensuring the reliability, availability, and lifecycle health of OpenAI’s compute infrastructure. We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments. As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses. About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses. This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability. You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems. Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts. You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments. The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills. Candidates must be able to sit onsite at our

AWSLinuxRestAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s Executive Briefing Center (EBC) Solutions Architect (SA) team is looking for a highly hands-on Solutions Architect with exemplary communication skills. The role involves developing, demonstrating (in the NVIDIA EBC), and packaging agentic AI systems. Partnering with account SAs you will co-develop proof of concepts (POC) and "uplift" their presentation quality to match the NVIDIA branding and messaging used with Executive meetings. This is a builder’s and presenter's role! You will spend time architecting and writing code. You will develop multi-agent systems, retrieval pipelines, and optimized inference stacks on NVIDIA’s full-stack accelerated computing platform. We want a creative, diligent, and curious engineer energized by agentic AI and ready to make significant change. If that’s you, join us! What you’ll be doing: Architect, build, and ship end-to-end Agentic AI applications for a variety of use cases—spanning multi-agent coordination, long-horizon reasoning, planning, and tool use. Act as Technical Advisor alongside fellow Subject Matter Experts (SME) in Executive Briefings. Creating and presenting demos that are used at Trade Shows or Customer Meetings. Partner with NVIDIA engineering, product, and sales teams to secure build wins, translate customer feedback into actionable product and roadmap insights, and scale global expertise through technical collateral, workshops, and developer communities. What we need to see: BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, AI/ML, or a related field (or equivalent experience) 2+ years as an ML/Software Engineer or Solutions Architect writing production-level code in Python and/or C/C++ in Linux environments. Validated experience building sophisticated agentic and multi-agent AI sy

PythonAWSAzureGCP
C
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -94.7%
Quick readStrong listing-quality and freshness signals

As an Engineering Manager on Coder’s Core Workspaces team, you’ll lead engineers building and evolving the systems behind our agentic development experience. You’ll help make agents more capable, reliable, and useful across real development environments. You’ll guide technical direction while growing the team and keeping execution sharp. You’ll work closely with Engineering, Product, and Design across the agent harness, integrations, and developer workflows. What you’ll do here Lead and grow a team within our Workspaces organization. Set technical direction across the agent harness, integrations, and workflows. Stay close to the code and contribute to architecture and implementation decisions. Evolve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve reliability, performance, and operability across agentic systems. Coach engineers, raise the technical bar, and create clarity around priorities and tradeoffs. What we’re looking for Experience managing and growing software engineering teams. Strong hands-on engineering experience with React and TypeScript . Experience with Go . Hands-on experience building systems around LLMs and agentic workflows. Experience with model APIs, tool calling, context management, or agent loops. Strong distributed systems knowledge. Working knowledge of AWS . Strong technical judgment and comfort working through ambiguity. A track record of helping engineers grow while maintaining a high execution bar. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP , agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute. Experience building abstractions across multiple model providers. Deep experience with AWS, Kube

TypeScriptReactAWSDocker
C
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -94.7%
Quick readStrong listing-quality and freshness signals

As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. To provide substantive overlap with the team, this position must be in Eastern Time. What you’ll do here Design and build production systems in Go, with work across React and TypeScript where needed. Improve agent execution, tool use, context management, streaming, and long-running workflows. Extend our provider-agnostic architecture as models and capabilities change. Build reliable integrations between agents, workspaces, tools, and developer infrastructure. Own projects from implementation through rollout and iteration. Contribute to design reviews, code reviews, and technical discussions. Partner with Product and Design to turn agent capabilities into useful developer experiences. Improve the reliability, performance, and operability of agentic systems. What we’re looking for Strong experience building and operating production software systems. Hands-on experience with Go. Experience with React and TypeScript. Experience building systems around LLMs or agentic workflows. Familiarity with model APIs, tool calling, context management, or agent loops. Good understanding of distributed systems and production reliability. Working knowledge of AWS. Strong problem-solving skills and comfort working through technical ambiguity. Someone who contributes beyond their own code through reviews, collaboration, and knowledge sharing. Bonus tacos if you have Experience building coding agents, developer tools, or cloud development environments. Experience with MCP, agent tools, or multi-agent systems. Experience with remote execution, sandboxing, or isolated compute.

TypeScriptReactAWSDocker
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

PythonKubernetesLinuxArtificial Intelligence
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $128K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor

PythonKubernetesLinuxAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
L
📍 Bethesda, United States
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

Leidos has an exciting opportunity for Cyber Security Engineer—Technical Lead in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This role is responsible for protecting the customer’s information systems and networks from potential cyber-attacks. The Cyber Security Engineer– Technical Lead will serve in a hands-on “player-coach" capacity, dedicating approximately 75% of time to direct technical engineering, troubleshooting, and implementation work, while providing technical leadership and coordination across the security team. The candidate must display an excellent understanding of technology and utilization of Firewalls (Security Groups), VPNs, Data Loss Prevention (DPS), IDS/IPS, Web-Proxy, Security tools, and Security Audits. Candidate will work directly with Team leads, developers, operations personnel, and other Technical Leads throughout a DevSecOps life cycle both on policy and technical implementation of technologies. This is not a supervisory management role. Success in this position is measured by individual technical contribution and resolution of complex security issues, in addition to technical leadership impact. Primary Responsibilities: Plan, implement, manage, monitor, and upgrade security controls and tools used to protect enterprise systems and networks, while identifying opportunitie

PythonReactAWSLinux
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Staff Engineer in Micron’s HBM Product Engineering Component Validation team, you will be an individual contributor responsible for the validation of advanced High Bandwidth Memory (HBM) products. You will use proven knowledge in dynamic random-access memory and high-bandwidth memory architecture to compose validation coverage, implement and analyze validation runs, lead product debug and failure analysis toward a zero-escape validation goal across test modes and product achievements. You will also adopt AI-enabled ways of working, using enterprise AI tools to accelerate your engineering tasks and improve efficiency, work within Micron’s global validation organization and collaborate with groups at different locations. Responsibilities: Translate DRAM/HBM develop intent and architecture into effective verification scope and test cases. Contribute to validation of new and high-risk product features, working with Design Validation, and Systems Engineering. Review design specifications, schematics, and datasheets, provide design feedback and represent the Product Engineering perspective in Design and Verification reviews. Read, analyze, and debug circuit blueprints to detect design-related failure mechanisms, and recommend approaches for verification and fault diagnosis. Perform root-cause analysis and examination of electrical failures (debug) of validation findings, distinguishing developed marginality from process and test-related issues. Apply automation and modern tooling (including

PythonGitLinuxAI
C
📍 Work From Hom, United States
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Designs, builds, and maintains large-scale data infrastructure and data processing systems. Implements robust and scalable solutions to support data-driven applications, analytics, and business intelligence. What you will do Ensures seamless integration of data from different sources, such as databases, application programming interfaces (APIs), or streaming platforms. Optimizes data processing and query performance by fine-tuning data pipelines, database configurations, and data partitioning strategies. Establishes data quality checks and validations to identify and resolve data issues, ensuring high-quality and reliable data for downstream applications and analytics. Implements security measures to protect sensitive data throughout the data lifecycle by working closely with security teams to ensure data encryption, access controls, and compliance with data protection regulations. Collaborates with cross-functional teams, including data scientists, analysts, software engineers, and business stakeholders. Designs and develops data infrastructure, including data warehouses, data lakes, and data pipelines. Establishes auditing and monitoring mechanisms to track data access and maintain data governance standards. Establishes monitoring and alerting mechanisms

PythonSQLGCPLinux
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As a Principal Engineer at Micron’s HBM New Product Validation team in Boise, ID you will be a senior technical guide responsible for validation strategy and readiness for future-generation High Bandwidth Memory (HBM) products. You will apply deep expertise in complex RAM and high-bandwidth memory architecture and building to establish architecture validation strategies. You will impact decisions that improve validation observability and debug efficiency . You will lead investigations of complex architecture, building , and silicon issues from pre-silicon simulation through post-silicon characterization. You will partner across Design Engineering, Design Validation, Product Engineering, Systems teams, and will develop and promote AI-enabled validation methodologies that improve efficiency, quality, and scalability. Responsibilities: Define architecture-aware validation strategies and coverage direction for future-generation HBM products. Drive validation readiness for next-generation products by preparing strategies, test environments, and execution plans before first silicon, reducing risk, accelerating issue resolution, and preventing late-stage surprises. Evaluate new DRAM/HBM architectural features and define validation requirements and testability needs before DBR. <span style="color:#

PythonGitLinuxAI
Z
📍 Bellevue, Washington, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h

PythonKubernetesLinuxAI
C-
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$175K – $215K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We’re looking for a Senior Software Engineer to join our Network Engineering team to accelerate building and scaling our innovative systems that support our growing identity platform. In this role, you will build the next-generation infrastructure that underpins all systems at CLEAR. The ideal candidate for this role will approach challenges with an eye toward reliability, simplicity, and scalability. What You'll Do: Develop and maintain a streamlined process for engineers to effortlessly build and deploy scalable and reliable software-defined networking solutions on AWS. Enhance our compute platform (Kubernetes) by integrating new functionalities and features, focusing on AWS networking services and concepts such as VPCs, Route Tables, Security Groups (SGs), ALBs/ELBs, and Route53, as well as implementing Kubernetes networking solutions like service mesh (Istio) to optimize service communication and management. Collaborate across engineering teams to advocate for and implement best practices in observability, utilizing tools like Splunk or Datadog to ensure robust network monitoring. Act as a product owner for our infrastructure, collecting feedback and requirements from engineering teams to address pain points and develop solutions, particularly in the realm of AWS networking and cloud-native design principles. What you're great at: 6+ years of extensive experience in infrastructure and platform development, particularly in software-defined networking and AWS cloud services. Proficient in writing production-grade softwar

PythonAWSKubernetesGit
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA has a rapidly expanding ecosystem of data center platform designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. Each brings together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will be instrumental in ensuring our data center solutions deliver industry-leading performance for accelerated computing workloads. What you will be doing: Design and execute comprehensive performance benchmarking strategies for our data center platforms and products Characterize real-world AI training, inference, and HPC workloads at scale Define, track, and report key performance indicators (throughput, latency, efficiency, scaling) Build automation tools and frameworks for performance monitoring and analysis Identify and analyze performance bottlenecks across compute, memory, network and storage subsystems Work closely with architecture, hardware,

PythonDockerKubernetesLinux
🔔

Get new linux system administrator jobs in United States by email

Daily job updates · Unsubscribe anytime