Jobs in United States

Software Engineer Infrastructure in United States

2,095 active opportunities · Updated October 2026

Explore current software engineer infrastructure jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

Hiring demand

51/100

steady · 562 related jobs

Hiring trend

-80.2%

Job postings compared with the previous 30 days

Remote options

15.8%

Share of matching jobs listed as remote

Typical salary

$177.2K – $177.2K/yr

Based on 32 salary observations

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
T
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to validate the System Management Controller (SMC) and enable seamless multi-chip integration. In this role, you will design and execute tests, build infrastructure, and debug issues across chiplet-based SoCs. You’ll have the opportunity to work with remote mentorship while contributing to the foundation of scalable multi-die systems. This role is hybrid, based out of Toronto, Ontario, Boston, MA or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proficient in SystemVerilog, SV-UVM, Python, and C/C++ with strong verification skills. Experienced in writing test plans, building infrastructure, and debugging hardware/software flows. Comfortable working with remote mentorship and distributed teams. Familiar with AI-assisted tools like Copilot, Cursor, and Claude to accelerate verification. What We Need Develop and maintain SMC tests and supporting DV infrastructure. Write, execute, and track test plans for chiplet and multi-chip SoC designs. Use C/C++ to develop tests compiled, loaded, and executed directly on the DUT. Triage, analyze, and debug issues in clos

PythonAWSAIC++
L
📍 Bethesda, Moldova, United States
✓ Quality checkedCompany trend +400%

Leidos has an exciting opportunity for a Sr. DevOps Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. Scripting and

JavaScriptPythonJavaAWS
V
📍 United States· Full-time
✓ Quality checkedCompany trend -94%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Core Platform team provides the foundational infrastructure that powers all engineering at Vanta. We're expanding upmarket to support enterprise customers, which requires strategic investment in platform systems that ensure security, reliability, and developer productivity at scale. As we expand upmarket to support enterprise and regulated customers, we’re investing heavily in platform capabilities that scale securely while reducing cognitive load for product teams. As the Engineering Manager, Core Platform at Vanta, you'll own the foundational infrastructure that every engineer builds on, ensuring it scales with company growth while remaining fast, simple, and reliable. This team’s ownership spans shared services infrastructure, observability and monitoring, datastore management, and async work systems. Our Engineering Managers develop and grow high-performing teams that deliver significant value to our customers and enable our business to scale. This role sits at the intersection of technical architecture and team development, with real authority to set direction and grow a world-class platform team. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an Engineering Manager at Vanta: Lead and grow high-performing platform engineerin

MongoDBAWSRestAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Engineering Manager, Communications, you'll lead the team responsible for in-game text chat — one of the most-used surfaces on Roblox and a cornerstone of how our community connects. Every day this system carries billions of messages across hundreds of millions of users, in real time, across 2D and 3D spaces, on every device we support. You'll own the roadmap and the engineering org behind it: building rich, immersive, and engaging communication experiences that let people and creators express themselves safely and seamlessly — better than in real life. This is a role for a leader who thinks like a product builder as much as an engineer. You'll balance the demands of massive scale and rock-solid reliability with a relentless focus on the user experience, shipping features that make conversation on Roblox feel effortless, expressive, and safe. You'll grow and develop a team of engineers, set technical direction, and partner deeply across product, design, trust & safety, and infrastructure to define what communication on Roblox becomes next. You Will Lead, grow, and develop a team of software engineers building the in-game text chat platform, setting a high bar for engineering

AWSGitAIGo
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -97.2%

From $127K/yr

Quick readStrong listing-quality and freshness signals

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

MongoDBAWSAzureGCP
I
📍 Oregon, Hillsboro, United States
✓ Quality checkedCompany trend +116%

Job Details: Job Description: The Role As a Cloud Application Development Engineer within Intel Manufacturing Foundry Cloud Services (imFCS) , you will design, develop, deploy, and support cloud-native applications that power semiconductor manufacturing, engineering automation, and AI-driven factory operations. You will build scalable, secure, and resilient platforms that improve engineering productivity, enable advanced analytics, and accelerate Intel Foundry's digital transformation. Key Responsibilities Develop and maintain cloud-native applications and services supporting manufacturing and engineering workflows. Support 24x7 manufacturing operations through on-call rotations, incident response, and root-cause analysis. Design and implement scalable microservices, APIs, and containerization systems utilizing Kubernetes orchestrator and cloud native applications. Architect secure cloud solutions spanning application, data, networking, identity, and observability domains. Build and maintain CI/CD pipelines, Infrastructure-as-Code, and DevOps automation. Provide technical leadership for contractor and partner development teams across multiple geographic regions, driving architecture, implementation, and operational excellence for cloud-native solutions. Collaborate with manufacturing, engineering, and software teams to deliv

PythonDockerKubernetesAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonAIExcelHR
B
📍 Berkeley, Macau S.a.r., United States
✓ Quality checkedCompany trend -20.1%

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

AWSAzureDockerKubernetes
L
📍 Hampton, Vatican City State (holy See), United States
✓ Quality checkedCompany trend +400%

The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. ​ Roles and Responsibilities: ​Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. ​Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. ​Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. ​Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. ​Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. ​Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. ​Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti

PythonAnsible
B
📍 Berkeley, Macau S.a.r., United States
✓ Quality checkedDemand 51/100Company trend -20.1%

$177.2K – $211K/yr · Jobiba est.

Associate and Mid-Level Software Engineers Company: The Boeing Company The Boeing Company is looking for an Associate or Mid-Level Software Engineer to join our team in Berkeley, MO. Position Responsibilities: Lab Environment Provisioning: Design, provision, and maintain scalable lab environments on-premises and on cloud platforms (Azure, AWS) using Infrastructure as Code (IaC) tools such as Terraform and Ansible. Automated Integration Testing Orchestration: Develop and manage automated integration testing workflows that aggregate inputs from multiple program segments, ensuring comprehensive test coverage and timely feedback. Deployment Management: Manage and optimize automated deployment pipelines and mechanisms for both physical and virtual systems, ensuring reliable and repeatable software delivery. Data-Driven Feedback & Reporting: Orchestrate processes to collect, analyze, and deliver comprehensive feedback on product stability, performance, and integration issues to segment development teams, enabling continuous improvement. Collaboration & Communication: Work closely with cross-functional teams including development, QA, security, and physical lab team to align deployment strategies, testing requirements, and environment configurations. Security & Compliance: Integrate security best practices into provisioning, deploying, and testing processes, supporting compliance with relevant standards and frameworks. This position is expected to be 100% onsite. The selected candidate will be required to work onsite in Berkeley, MO. Basic Qualifications (Required Skills/ Experien

PythonAWSAzureDocker
B
📍 Berkeley, Macau S.a.r., United States
✓ Quality checkedDemand 51/100Company trend -20.1%

$177.2K – $211K/yr · Jobiba est.

Flight Simulation Software Engineers (Associate, Experienced, and Senior) Company: The Boeing Company The Boeing Company is currently seeking Flight Simulation Software Engineers (Associate, Experienced, and Senior) to join the Flight Simulation Labs team located in Berkeley, MO . This position will focus on supporting the Boeing Defense, Space & Security (BDS) business organization. **Our teams are currently hiring for a broad range of experience levels including; Associate, Experienced, and Senior Level Software Engineers. Our Air Dominance software engineers design, develop, and demonstrate avionics solutions for the platforms we build and upgrade. Your efforts will assist in integration of innovative technologies to maintain our platform’s dominance over the battlefield for many decades to come. Additionally, our software engineers support the Flight Simulation products associated with our platforms. The roles will involve software development for tools that support simulation infrastructure across Flight Simulation Labs. Selected candidates will be aligned to a statement of work that best suits their interests and skillset, some of which could include: Creating core simulation infrastructure to serve as a jumping-off point for programs developing flight simulation capabilities Developing tools to simulate, monitor, and capture messages between subsystems in a lab environment Modeling various aircraft capabilities, including the ability to run constructive scenarios in a testing environment Position Responsibilities: Develops, documents and maintains architecture, requirements, algorithms, interfaces and designs for softwa

PythonJavaC#Recruitment
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

AIGoExcelSEM
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -97.2%

From $151K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for a Senior Engineering Manager who is ready to lead through ambiguity and improve how software gets built at MongoDB. This role leads teams focused on developer productivity, with an emphasis on measurable improvements to the software development lifecycle. This role can be based remotely in the United States. The Team The AXIS team (AI, X-functional tools, Insights, and Signals) sits within Developer Productivity and is responsible for overseeing the metrics and observability infrastructure of our expansive developer environment to help build a strong data-driven culture. You’ll also be a key partner in building the agentic ecosystem for AI-driven development across engineering. Candidate Profile We’re looking for an experienced leader with a passion for solving the big challenge of measuring developer productivity and providing the actionable signals that help teams improve their performance. They should be comfortable working collaboratively with other leaders and partners across our Engineering and Data teams in maximizing the use of data for insights and AI enablement. The right candidate for this role will have 4+ years of experience managing software engineers, including hiring, performance management, growth planning, and compensation; required for external candidates and preferred for internal candidates 8+ years of hands-on software engineering experience building and operating production systems; experience in developer tooling, platform engineering, observability, or data engineering is a strong plus Demonstrated the ability to lead through ambiguity, work across team boundaries, and deliver outcomes without close supervision Strong customer orientation and sound judgment in finding practical, high-leverage solutions Experience working with systems involving analytics, data pipelines, and metrics platforms Experience with AI tools development and enablement efforts Strong technical judgment, including the ability to evaluate t

MongoDBAWSAzureAI
S
📍 Bellevue, WA, United States· Full-time
✓ High-confidence listingCompany trend -92%
Quick readStrong listing-quality and freshness signals

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day. The Grid Service and Platform Engineering team is looking for a highly motivated and collaborative Software Engineering Manager. This role involves leading the engineering of mission-critical, tier 0 service infrastructure, the foundational data platform that powers Smartsheet at scale. You will oversee services that handle millions requests per day, operate at 99.999% availability, and deliver low-latency, high-throughput performance for millions of customers worldwide. We are an agile team that operates iteratively, focused on building high-quality software and adhering to rigorous operational best practices across complex, cross-functional distributed systems. This full-time position reports to the Director, Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. You Will: Manage one or more related teams of 6–10+ software engineers, driving development of tier 0 grid services and platform infrastructure that millions of customers depend on daily. Own and uphold 99.999% service availability targets across critical platform services, embedding reliability engineering, incident management, and on-call rigor into team culture. Help architect and guide technical vision to evolve low-latency, high-throughput service platforms capable of sustaining millions requests per day with predictable, consistent performance under load. Guide and mentor engineers on distributed systems architecture, scalability patterns, and platform best pr

VueAWSAgileScrum

Related career options

Similar roles with stronger pay

Client Service Associate

Demand 46/100 · 8 jobs

$840K – $840K/yr

Salary →
Director of Product

Demand 43/100 · 6 jobs

$382.5K – $382.5K/yr

Salary →
Physical Design Engineer

Demand 43/100 · 8 jobs

$300K – $300K/yr

Salary →
Sr. Engineer

Demand 42/100 · 7 jobs

$300K – $300K/yr

Salary →
Senior Director

Demand 38/100 · 30 jobs

$278.9K – $278.9K/yr

Salary →
Senior Product Designer

Demand 30/100 · 11 jobs

$255.7K – $255.7K/yr

Salary →
🔔

Get new software engineer infrastructure jobs in United States by email

Daily job updates · Unsubscribe anytime