Jobs in United States

Linux System Administrator in United States

217 active opportunities · Updated October 2026

Explore current linux system administrator jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $326.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer, you will be reporting to the Manager of Application Security leading the Security Design and Review pod. You will play a pivotal role in shaping the growth of Information Security's (InfoSec) Application Security team, collaborating with engineering teams early in their processes to provide secure design solutions and establish security standards. Your responsibilities will include threat modeling, secure system design, automation, and penetration testing. As a key member of the team, you will drive company-wide projects across diverse tech stacks, working with engineering leaders to remediate security challenges. You will define and evolve the technical vision for scaling application security practices across the organization. You will: Lead company-wide security initiatives to address critical security challenges. Build and nurture cross-company relationships to achieve security objectives. Provide guidance on product security processes and standards. Define and expand partnerships with key engineering teams across Roblox. Apply critical thinking and analytical skills to develop security protocols and communicate effectively with stakeholders. Re

PythonAWSGitLinux
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -84.7%

From $244K/yr

Quick readStrong listing-quality and freshness signals

Role Summary: Datadog is seeking a Staff Software Engineer to help shape the future of our Bring Your Own Cloud (BYOC) Logs offering by unifying observability pipelines with log management software that customers deploy and manage in their own infrastructure. This role will focus on building and scaling systems that process, route, and store high-volume observability data within customer-managed infrastructure. You will operate as a hands-on technical leader, driving architecture, cross-team delivery, and product direction across a complex and evolving space. This is a high-impact opportunity to influence product strategy, mentor engineers, and solve deeply technical challenges at scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Make customer-controlled deployments feel like a managed Datadog product: deployment, upgrades, configuration, observability, diagnostics, reliability, and secure operation across diverse customer cloud environments Build and scale high-throughput systems for log processing, routing, and transformation across distributed environments Lead cross-team initiatives, aligning engineers, product managers, and stakeholders to deliver complex, multi-team projects Design and implement software that runs reliably that customers deploy and operate within their own cloud infrastructure. Improve system performance, scalability, and cost efficiency through thoughtful trade-off analysis and capacity planning Contribute hands-on to critical code paths, debugging, and deployment challenges in customer environments Who You Are: You have significant experience building software that is installed, deployed, and operated in customer environments rather than only as a fully managed SaaS service. You have strong expertise in distributed systems,

AWSAzureGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,

AWSKubernetesCI/CDLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

PythonSQLAWSLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati

PythonSQLAWSLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet Hardware team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automation to reduce manual work

PythonSQLAWSLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics, and vector database that powers all vector search and retrieval augmented generation (RAG) at OpenAI. In addition to retrieval, as an online database, Rockset powers core functionality across all of OpenAI's product lines and many critical internal use cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal performance optimization (our core engine is written in C++), and building scalable database infrastructure from the ground up. As an engineer on the Database Systems team, you'll contribute to the core database engine, driving improvements across ingestion, query execution, indexing, and storage. You'll partner with teams across OpenAI to unlock new product capabilities and help scale online database reliability and throughput as usage grows by orders of magnitude. In this role you will: Design, build, and operate high-performance distributed systems Identify and resolve performance bottlenecks to scale infrastructure to the next order of magnitude Define long-term technical direction and guide system evolution Collaborate with product, engineering, and research teams to deliver scalable and reliable infrastructure Dig deep into complex production issues across the stack Contribute to incident response, postmortems, and best practices for system reliability You might thrive in this role if you: Have significant experience building, scaling, and optimizing distributed systems at scale Are curious about database internals, storage engines, or low-latency query systems Enjoy debugging challenging performance issues in complex, high-throughput systems Have experience operating production clusters at scale (e.g., Kubernetes or other orchestration systems) Think rigorously about scalability, correctness, and reliability Thrive in fast-paced environments with high

AWSAzureGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning

PythonAWSCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil

PythonAWSCI/CDGit
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -92.9%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We’re hiring talented Software Engineers to join the Snowflake Database Engineering team! In this role you will work directly with our team to evolve our elastic, large scale, high-performance data processing system. We need smart engineers who can pick up and understand complex technical areas quickly – and who are enthusiastic about building new technologies! AS A SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: Design, develop, and support a petabyte-scale cloud database that is highly parallel and fault-tolerant. Build high-quality and highly reliable software to meet the needs of some of the largest companies on the planet. Analyze and understand performance and scalability bottlenecks in the system and solve them. Pinpoint problems, instrument relevant components as needed, and ultimately implement solutions. Design and implement novel query optimization or distributed data processing algorithms which allow Snowflake to provide industry leading data warehousing capabilities. Design and implement the new service architecture required to enable the Snowflake Data Cloud Develop tools for improving our customers' insights into their workloads. OUR IDEAL SOFTWARE ENGINEER WILL HAVE: 2+ years industry experience working on commercial or open-source software. Fluency in Java or C++

JavaVueSQLPostgreSQL
T
📍 Austin, Texas, United States
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Our diverse team of technologists have developed a high performance RISC-V-based CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance teams. This is the plumbing that both teams' performance work stands on — benchmarking automation, data collection, and workload capture across silicon, FPGA/emulation platforms and performance models. You'll work across both teams, enabling performance and software engineers to spend their time on analysis and optimization instead of running experiments by hand. This role is hybrid, based out of Austin, TX or Santa Clara, CA. We welcome candidates at various experience levels. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Great at identifying problems and developing solutions, with a bias toward owning the systems you build. Enjoys building tools and optimizing workflows so other engineers can move faster. Strong Linux systems engineer, c

PythonLinuxAI
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -85.9%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About Replit Replit is building the world’s most ubiquitous AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it’s an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit is also the leader in secure vibe coding. We protect apps, give users features to manage security risks, and help them vibe code more safely. About the Role In this role you will build powerful tools that help product engineers iterate rapidly on the Agent experience and directly enhance the core Agent itself. You’ll bridge the gap between the AI team (working on the core Agent logic) and the UX team (crafting delightful Agent experiences), enabling both groups to excel within their specialties. This role blends systems engineering, developer experience and product engineering. We tackle complex challenges across the full stack, from browser-based interfaces to high-performance backends to Linux systems engineering. We’re looking for engineers who have a keen sense of the product experience and how to power it with performant systems. On this team, you’ll have the opportunity to grow your skills across our infrastructure and product, and to lead end-to-end efforts with meaningful impact. We value diverse perspectives and encourage candidates from all backgrounds and experiences to apply. You Will Build high-throughput backend applications and services, like streaming chat between user and agent. Design a collaborative "Multiplayer Computer" that lets humans and AI agents work together on shared shells, filesystems, and state—conflict-free and in real time. Develop infrastructure (frontend & backend) that empowers product enginee

GitLinuxAI
G
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -97.9%

From $126.4K/yr

Quick readStrong listing-quality and freshness signals

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As an Assigned Support Engineer, you’ll be a trusted technical advisor to GitLab’s largest Self-managed, GitLab Dedicated, and GitLab.com customers, helping them avoid operational disruption and get the most from GitLab. You’ll combine deep Linux systems expertise, GitLab and CI/CD knowledge, and a proactive support mindset to understand each customer’s environment, anticipate and prevent issues, and solve complex technical and business challenges. In a typical week, you might be prioritizing strategic blockers with customer stakeholders, partnering with Product, Development, Infrastructure, Customer Succ

KubernetesCI/CDGitLinux
L
📍 San Diego, California, Canada
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

The Undersea Systems Division at Leidos currently has an opening for a Senior Systems Engineer to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Senior Systems Engineer and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, action-item logs, a

LinuxProject ManagementPmp
L
📍 San Diego, California, Canada
✓ High-confidence listingCompany trend +500%
Quick readStrong listing-quality and freshness signals

The Undersea Systems Division at Leidos currently has an opening for a Systems Engineer to support a customer site in San Diego, California. This is an exciting opportunity to apply systems engineering, naval operations, test and evaluation, and technical project-management experience in support of large-scale fleet activities and critical maritime missions. The selected candidate will serve as the onsite Systems Engineer and senior technical representative, working closely with the Leidos Project Manager, customer leadership, engineers, fleet personnel, operational stakeholders, and government representatives. The position will support day-to-day onsite execution of program activities while also contributing directly to systems engineering, integration, test, and evaluation efforts. Primary Responsibilities: Serve as the onsite senior technical representative for Leidos at the customer site in San Diego. Work side by side with customer leadership and technical personnel to support day-to-day program execution and mission priorities. Support the Leidos Project Manager with program planning, execution, customer coordination, schedule management, risk management, and delivery of program commitments. Lead and coordinate engineering activities across multiple technical disciplines and stakeholder organizations. Translate customer operational needs into technically sound engineering approaches, priorities, requirements, and recommendations. Coordinate with engineers, subcontractors, fleet personnel, operators, government stakeholders, and program leadership to resolve technical and operational issues. Track technical scope, milestones, schedules, deliverables, action items, risks, issues, dependencies, and decisions. Maintain project-management artifacts, including project plans, schedules, milestone trackers, risk and issue registers, action-item logs, and status rep

LinuxProject ManagementPmp
🔔

Get new linux system administrator jobs in United States by email

Daily job updates · Unsubscribe anytime