Jobs in United States

Infrastructure Engineer in United States

1,475 active opportunities · Updated October 2026

Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

PythonAWSAzureCI/CD
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel

PythonKubernetesLinuxArtificial Intelligence
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore fosters continuous learning and innovation. Job Summary Reporting into the Systems Engineering organisation, the Distinguished Engineer, End-to-End Security Architect will define and lead the security architecture for Graphcore’s inference service platform. This role is responsible for establishing a comprehensive security strategy spanning platform, infrastructure, networking, service operations, customer assurance, and compliance readiness. Working across multiple engineering and operational functions, the successful candidate will provide technical leadership, drive security requirements, and ensure the platform delivers robust protection, resilience, and trust for customers. The Team You will work closely with teams across security architecture, infrastructure engineering, networking, site reliability engineering, platform software, firmware, data centre operations, compliance, legal, customer engineering, and customer security. The team collaborates across the business to deliver secure, reliable, and scalable AI infrastructure and services while supporting customer assurance, regulatory requirements, and operational excellence. Responsibilities and Duties Own the end-to-end security a

AIRustExcelRecruitment
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -72.4%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa

PythonSQLPostgreSQLRedis
E(
📍 San Francisco Bay Area, California, United States· Full-time
✓ High-confidence listingCompany trend -100%

$135K – $225K/yr

Quick readStrong listing-quality and freshness signals

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. Who you are We are seeking an experienced DevOps Engineer to join our growing team and play a pivotal role in designing and building our platform and infrastructure as we continue to scale our product and user base. As a part of our team, you will be working in a dynamic, fast-paced environment to ensure the reliability, scalability, and performance of our systems, while focusing on service architecture and deployment, query optimization, distributed systems, data and machine learning infrastructure, and security and authentication. Most importantly, you are excited to be part of a mission-oriented, fast-paced, high-growth startup that can create a lasting impact. You will: Partner with product teams to architect, design, and build the foundational infrastructure for our products. Design, develop, and deploy highly available and scalable Multi-tenant SaaS solutions on any one of the public cloud networks like AWS, Azure and GCP. Leverage technologies such as Kubernetes, Helm, Terraform, and Istio to achieve infrastructure resilience. Drive the automation of infrastructure tasks, from provisioning to configuration management and deployment, utilizing tools like Terraform, Ansible, a

AWSAzureGCPKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team Security is at the foundation of OpenAI's mission to ensure that artificial general intelligence benefits all of humanity. The Identity Infrastructure Engineering team sits at the core of this effort, designing and building the identity and access management solutions that protect model weights, customer data, and critical systems across multiple cloud environments. The team partners across OpenAI, including Applied Engineering, Research, IT, Security, Infrastructure, and Engineering, to provide secure and scalable platforms for identity, access management, permissioning, orchestration, and safe AI research. About the Role We’re looking for an engineering leader to lead Identity Infrastructure Engineering, the team building the systems that govern and scale access across OpenAI’s research, engineering, and internal platforms. This role sits at the center of cloud infrastructure, identity, software engineering, and security-critical operations. You’ll lead engineers building control planes, policy systems, workload and agent authorization patterns, infrastructure-as-code, and operational foundations that help OpenAI move quickly while keeping access reliable, auditable, least-privileged, and safe under failure. The ideal candidate has led teams responsible for large-scale, mission-critical infrastructure. They can go deep into code and architecture when needed, while giving engineers and technical leads the clarity and ownership to do their best work. They set technical direction, grow strong teams, make durable architecture decisions, and turn ambiguous 0-to-1 problems into platforms OpenAI can trust and build on for years. In this role, you will: Build and lead a high-performing Identity Infrastructure team, going deep enough technically to set direction while empowering the team to own delivery. Define the strategy for identity platform as the policy plane for access across people, agents, workloads, services, clouds, and internal systems. Scale Acc

AWSGitRestAI
C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is seeking a highly skilled Staff Data Engineer, Observability Engineering to join the Enterprise Observability Platform organization and help advance the next generation of observability, infrastructure, and security data capabilities. The Staff Data Engineer, Observability Engineering will play a critical role in designing, building, and operating scalable data pipelines and data products that power enterprise observability, operational intelligence, and security analytics across the organization. The Staff Data Engineer, Observability Engineering is a senior individual contributor responsible for developing and optimizing Databricks-based data engineering solutions that ingest, transform, govern, and deliver high-volume telemetry, infrastructure, application, and security data. This role combines deep hands-on technical execution with ownership of engineering excellence, operational reliability, performance optimization, and data platform best practices. Working closely with Observability Engineering, Security Engineering, Infrastructure Engineering, and Data Platform teams, the Staff Data Engineer, Observability Engineering will contribute to the evolution of the enterprise observability lakehouse by building resilient ingestion frameworks, establishing data quality standards, enhancing governance controls, and driving efficient, scalable data processing patterns. The id

PythonSQLAzure
C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is seeking a highly skilled Senior Data Engineer, Observability Engineering to join the Enterprise Observability Platform organization and help advance the next generation of observability, infrastructure, and security data capabilities. The Senior Data Engineer, Observability Engineering will play a critical role in designing, building, and operating scalable data pipelines and data products that power enterprise observability, operational intelligence, and security analytics across the organization. The Senior Data Engineer, Observability Engineering is a senior individual contributor responsible for developing and optimizing Databricks-based data engineering solutions that ingest, transform, govern, and deliver high-volume telemetry, infrastructure, application, and security data. This role combines deep hands-on technical execution with ownership of engineering excellence, operational reliability, performance optimization, and data platform best practices. Working closely with Observability Engineering, Security Engineering, Infrastructure Engineering, and Data Platform teams, the Senior Data Engineer, Observability Engineering will contribute to the evolution of the enterprise observability lakehouse by building resilient ingestion frameworks, establishing data quality standards, enhancing governance controls, and driving efficient, scalable data processing patter

PythonSQLAzure
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%
Quick readStrong listing-quality and freshness signals

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

AWSRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are seeking a senior data infrastructure engineer to join our new Data Insights team. Our team owns the data tooling that empowers Roblox builders to independently make informed and timely data-driven decisions. As an engineer on the team, you’ll work on the platforms behind tools like Superset, Hex, and Python notebooks, which provide critical insights into the health of our business to users at every level of the company. We tackle diverse challenges in data engineering, infrastructure, and analytics, to deliver the insights our customers need. You will collaborate closely with engineers across our data ecosystem to shape the future of product analytics at Roblox. This role offers the chance to be a founding team member and help define both the technical direction and the long-term shape of the product area from the ground up. This role is a great fit for you if you are proficient in designing and scale robust data infrastructure and applications and have a zeal for developing inspiring, easily maintainable, and reusable code. Join our team and make a significant impact at Roblox. You Will: Architect and deliver a high-pe

TypeScriptPythonReactSQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world. Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure. About the Role We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology. You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more. There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future. In this role, you will: Build, lead, and grow high-performing infrastructure engineering teams. Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search. Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach. Champion pragmatic use of agent technology to amplify execution velocity. Reduce operational toil and incident frequency through better abstractions, gua

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec

TypeScriptPythonJavaSQL
B
📍 Colorado Springs, United States
✓ High-confidence listingCompany trend +515.8%
Quick readStrong listing-quality and freshness signals

Senior DevSecOps Software Engineer Company: The Boeing Company The Boeing Company has an exciting opportunity for a Senior DevSecOps Software Engineer to support the Protected Tactical Enterprise Service (PTES) team in Colorado Springs, CO. As a DevSecOps Engineer on the Protected Tactical Enterprise Service (PTES) team, you will be responsible for integrating security practices within the development and operations lifecycle of tactical enterprise systems. You will collaborate closely with software developers, security teams, and infrastructure engineers to design, implement, and maintain secure, automated CI/CD pipelines, ensuring the delivery of resilient and compliant services in protected operational environments. Position Responsibilities: Leads the design, development, analyses, and maintenance of software systems that meet industry, customer and internal quality, safety, security and certification standards Partners with appropriate stakeholders to inform system definition and reviews translation of system-level requirements into software requirements and models that meet customer, operational and performance requirements and have clear traceability to design, code and test artifacts Reviews completion of software system-level analyses to identify risk, issues and opportunities; leads integration and deployment of mitigation actions throughout the software lifecycle. Leads code reviews to ensure alignment to requirements and standard Leads monitoring and reviewing test completion, verification processes and issue resolution for software systems Leads development of user documentation and training to educate end users about usage of software products Leads review of pr

PostgreSQLAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t

AWSRestAIRust
MT
📍 Boise, ID - Main Site, United States
✓ High-confidence listingCompany trend +1266.7%
Quick readStrong listing-quality and freshness signals

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As part of Micron's Technology Engineering & Innovation (TE&I) organization, you will have the opportunity to shape the future of our global infrastructure platforms while enabling business growth, operational resilience, and digital transformation at scale. The Opportunity Micron is seeking a transformational Senior Director of Technology Engineering & Innovation (TE&I) to lead the strategy, engineering, operations, and modernization of our global infrastructure ecosystem. This role is responsible for defining and executing Micron's vision across enterprise networks, cloud platforms, data center strategy, database services, infrastructure engineering, automation, observability, and global infrastructure operations. As a key member of the TE&I leadership team, you will drive innovation, operational excellence, and strategic transformation while building a high-performing organization focused on business outcomes and exceptional customer experiences. The successful candidate will be equally comfortable developing multi-year technology strategies, leading large-scale infrastructure transformations, driving operational performance, developing talent, and fostering a culture of collaboration, accountability, and continuous improvement. What You Will Do Lead Global Infrastructure Strategy & Transformation Define and execute Micron's global infrastructure vision and strategy. Develop multi-year roadmaps for: Enterprise Network Services <l

AIRecruitment
🔔

Get new infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime