Jobs in United States

Infrastructure Team Manager in United States

1,531 active opportunities · Updated October 2026

Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role OpenAI is seeking a Principal Security Engineer to join our Infrastructure Security (InfraSec) team. InfraSec protects the foundations of OpenAI’s research and production environments, spanning GPU supercomputing clusters, multi-cloud infrastructure, datacenters, networking, storage, and the critical services that power our frontier AI models. Our charter includes securing everything from bare-metal hardware and firmware, to Kubernetes clusters and service meshes, to data storage and access pathways for highly sensitive model weights and user data. As a principal engineer, you will set technical direction and drive execution on high-impact infrastructure security programs, partnering across various orgs at OpenAI to deliver durable controls that raise the security bar at OpenAI scale. In this role, you will: Own end-to-end security outcomes for one or more critical infrastructure areas, including multi-quarter strategy, roadmap, and delivery. Design and build security controls across diverse layers (e.g., physical hardware, firmware/BMC, OS, Kubernetes, networks, and CI/CD) to defend against sophisticated adversaries and insider threats. Lead cross-functional programs to deploy security enhancements and control changes across broad-scale infrastructure, balancing security guarantees with reliability and velocity. Take a generalist approach to building security controls, balancing a mix of security expertise and broad technical skillsets

AWSAzureKubernetesCI/CD
P
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.4%

From $139.8K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest brings millions of people the inspiration to create a life they love. Behind that experience is a complex infrastructure ecosystem that powers reliability, performance, measurement, and efficiency across the platform. As Pinterest grows, it’s increasingly important that we understand these systems clearly so we can make smarter decisions for both Pinners and the business. We’re looking for a Data Scientist to join our Infrastructure Data Science team. In this role, you’ll partner with engineering and cross-functional teams to make Pinterest’s infrastructure more measurable, intelligible, and actionable. Depending on the area, your work may span app performance, shopping infrastructure, metrics quality, infrastructure governance, or site reliability. You’ll help build the data foundations, measurement systems, and analytical fram

SQLAWSRestAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

Our work at NVIDIA is dedicated towards a computing model focused on visual and AI computing. For two decades, NVIDIA has pioneered visual computing, the art and science of computer graphics, with our invention of the GPU. The GPU has also shown to be spectacularly effective at solving some of the most complex problems in computer science. Today, NVIDIA’s GPU simulates human intelligence, running deep learning algorithms and acting as the brain of computers, robots and self-driving cars that can perceive and understand the world. We are looking to grow our company and teams with the smartest people in the world and there has never been a more exciting time to join our team! The AI Infrastructure Product Design team creates software used by engineers and researchers to prepare data, run complex workflows, and understand results. This internship offers ownership of a defined product problem from early research through a tested design and implementation handoff. Designers on this team often move between Figma and working HTML prototypes, and may hand off HTML directly to engineering. This work calls for a high standard of visual and interaction design alongside technical fluency. The role is a good fit for someone who enjoys making technically complex systems easier to understand and who uses large language models and software agents thoughtfully as part of their design and prototyping process. What you will be doing: Own a focused design project for an internal AI infrastructure product, from understanding the problem through a validated design and implementation handoff. Interview engineers and researchers, map their workflows, and turn the findings into clear product requirements, user flows, and interaction models. Create precise, implementation-ready interface designs and interactive prototypes in Figma and HTML/CSS, with careful attention to typography, hierarchy, spacing, visual consistency, interactio

JavaScriptReactAI
P
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.4%

From $114.3K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Pinterest brings millions of people the inspiration to create a life they love. Behind that experience is a complex infrastructure ecosystem that powers reliability, performance, measurement, and efficiency across the platform. As Pinterest grows, it’s increasingly important that we understand these systems clearly so we can make smarter decisions for both Pinners and the business. We’re looking for a Data Scientist to join our Infrastructure Data Science team. In this role, you’ll partner with engineering and cross-functional teams to make Pinterest’s infrastructure more measurable, intelligible, and actionable. Depending on the area, your work may span app performance, shopping infrastructure, metrics quality, infrastructure governance, or site reliability. You’ll help build the data foundations, measurement systems, and analytical fram

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Infrastructure Engineering function sits within IT and is responsible for reliably building, deploying, and operating critical on prem and hybrid environments that power internal services and critical R&D environments. This is an early, high-leverage technical role focused on applying strong Site Reliability Engineering discipline to environments where uptime, safety, recoverability, and security are non-negotiable. This person helps replace bespoke, one-off infrastructure with standardized infrastructure-as-code building blocks that compound reliability and operational leverage as OpenAI scales. About the Role We are looking for an experienced Site Reliability Engineer working on security infrastructure to design, build, and operate reliable, secure, and scalable infrastructure that underpins identity, access, endpoint, and shared platform services across the company. In this role, you will be a senior technical owner for infrastructure and identity systems end to end, from architecture and implementation through policy enforcement, upgrades, recovery, and day-two operations. You will build durable, production-grade platforms that remove operational friction, enforce security by default, and enable teams to move faster with confidence. This role is well suited for a hands-on senior engineer who thrives in ambiguity, enjoys owning complex systems end to end, and raises the reliability and security bar by replacing fragile implementations with standardized, repeatable infrastructure. This role is based in our San Francisco HQ and requires in-office presence. In this role, you will: Design, build, and operate reliable infrastructure across on-prem, hybrid, shared, and product adjacent environments. Establish standardized infrastructure patterns that replace bespoke implementations with repeatable, auditable, secure-by-default systems. Own the lifecycle of critical infrastructure platforms, including provisioning, deployment, upgrades, patching,

AWSAzureRestAgile
C
📍 Work At Home Arizona, United States
✓ Quality checkedCompany trend +340.2%

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: We are looking for a Staff Systems Engineer - Digital to join our team, the foundational layer that powers access, governance, and intelligence across our digital products. You will work horizontally across engineering, product, and AI teams, leading, owning, and evolving the shared infrastructure that every team at the company depends on. You will be the primary authority on how information is modeled, governed, and served across operational, analytical, and AI workloads - driving quality, compliance, and reliability at scale. If you thrive in a role where your architecture decisions multiply the productivity and capability of entire teams, this is the opportunity for you. Key Responsibilities: Data Architecture & Platform Ownership: Define and own the enterprise data architecture strategy across operational, analytical, and AI/ML workloads Design and govern data models, data contracts, and canonical schemas used across product and platform teams Evaluate and standardize data platform tooling — data lakes, warehouses, streaming, and serving layers (GCP BigQuery, Pub/Sub, Dataflow, or equivalent) Serve as the primary point of contact and SME for shared data platform concerns acro

PythonSQLAzureGCP
R
📍 Foster City, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -87.5%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Support Systems Lead at Replit, you'll build and own the systems that let support scale as fast as the product does. You'll own the tools our team works in every day, the technical setup behind self-service and AI-driven support, and the infrastructure that keeps all of it current as Replit ships. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You'll shape how our support systems adapt to new products, new surfaces, and AI-assisted workflows, operating effectively in ambiguity and turning ad hoc fixes into infrastructure the whole team can rely on. You'll combine hands-on technical depth with systems thinking to keep builders moving, whether they get unblocked through self-service, an AI agent, or a person. This is an individual contributor role to start, with room to grow and build out a team as support scales. IN THIS ROLE YOU WILL: Configure and maintain Zendesk and the surrounding support stack, from the day-to-day workflow and automation setup to business rules and permissions, keeping it able to flex and scale as needs change. Build and maintain assignment logic, queues, tagging and taxonomy, and escalation paths, keeping them running cleanly as volume and workflows change. Set up and maintain support tooling, workflows, and access across internal agents, outsourced vendors, and regions, keeping the systems working for each group as the stack changes. Partner with Engineering on the technical requirements for self-service and in-product support surfaces, and build the entry points, help widgets, and routing behind them. Spot the repetitive steps in agent and admin workflows before they become bottlenecks, and build the automations, bulk acti

PythonAISEMLean
R
📍 New York, New York, United States· Full-time· Remote
✓ Quality checkedCompany trend -87.5%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Support Systems Lead at Replit, you'll build and own the systems that let support scale as fast as the product does. You'll own the tools our team works in every day, the technical setup behind self-service and AI-driven support, and the infrastructure that keeps all of it current as Replit ships. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You'll shape how our support systems adapt to new products, new surfaces, and AI-assisted workflows, operating effectively in ambiguity and turning ad hoc fixes into infrastructure the whole team can rely on. You'll combine hands-on technical depth with systems thinking to keep builders moving, whether they get unblocked through self-service, an AI agent, or a person. This is an individual contributor role to start, with room to grow and build out a team as support scales. IN THIS ROLE YOU WILL: Configure and maintain Zendesk and the surrounding support stack, from the day-to-day workflow and automation setup to business rules and permissions, keeping it able to flex and scale as needs change. Build and maintain assignment logic, queues, tagging and taxonomy, and escalation paths, keeping them running cleanly as volume and workflows change. Set up and maintain support tooling, workflows, and access across internal agents, outsourced vendors, and regions, keeping the systems working for each group as the stack changes. Partner with Engineering on the technical requirements for self-service and in-product support surfaces, and build the entry points, help widgets, and routing behind them. Spot the repetitive steps in agent and admin workflows before they become bottlenecks, and build the automations, bulk acti

PythonAISEMLean
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are hiring a Sim Infrastructure Engineer to turn simulation systems into reliable, automated, production-quality pipelines that power model training, evaluation, and hardware-in-the-loop validation. This role owns the automation, orchestration, and tool integration that apply simulation to concrete robotics tasks: building CI/CD for SIL/HIL, presubmit checks, automatic model evaluation, metric computation and reporting, and the runtime infrastructure to run simulations at scale. You will collaborate closely with Sim Realism, Sim Environments, Research, and Ops to make simulation an integrated, reproducible, and measurable part of our ML and robotics workflows. This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role, you will: Build and maintain presubmit checks, continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible. Implement end-to-end automation to run model evaluation in sim (SIL) and orchestrate HIL runs; compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable. Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts, imitation-data collection, and presubmit model checks. Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), sol

PythonAWSKubernetesCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Industrial Compute team is responsible for building and scaling large-scale compute capacity across first-party data centers, strategic partners, and industrial infrastructure environments. We focus on converting power, land, hardware, and operational execution into reliable compute capacity that can support frontier AI training and inference workloads. This team operates at the intersection of infrastructure delivery, hardware systems, utilities, supply chain, and capacity strategy—ensuring OpenAI can scale compute faster than traditional models allow. About the Role We are seeking a Tokens-as-a-Service (TaaS) Lead to drive the end-to-end conversion of industrial-scale infrastructure investments into usable token capacity for OpenAI workloads. In this role, you will own execution across complex compute programs where raw infrastructure capacity must be transformed into operational GPU throughput. You will coordinate across data center delivery, power, networking, hardware deployment, workload enablement, finance, and external partners to ensure capacity becomes productive tokens as quickly and efficiently as possible. This role is ideal for someone who can bridge physical infrastructure delivery with compute utilization outcomes. Success requires strong systems thinking, elite program leadership, and the ability to drive accountability across internal teams and strategic partners. In this role, you will Lead Tokens-as-a-Service programs across industrial compute environments, including first-party and partner-owned capacity. Convert delivered power, space, and hardware capacity into production-ready token throughput. Build integrated execution plans spanning construction, power energization, rack deployment, networking, cluster readiness, and workload onboarding. Partner with infrastructure engineering, hardware, networking, finance, supply chain, and operations teams. Drive external providers, EPCs, OEMs, utilities, and strategic partners t

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

PythonAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet. Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions. With increasingly large supercomputers, the stakes continue to rise. Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale. This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Build and maintain automation systems for provisioning and managing server fleets. Develop tools to monitor server health, performance, and lifecycle events. Collaborate with clusters, networking, and infrastructure teams. Partner with external operators to ensure a high level of quality. Identify and fix performance bottlenecks and inefficiencies. Continuously improve automati

PythonSQLAWSLinux
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -84.1%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We're seeking a Software Engineer to join our First-Party Hardware team. In this role, you will design, build, integrate, and validate the software used to manufacture, qualify, and deliver our hardware from the factory. You will work across the stack to create the infrastructure that runs internally and externally to coordinate all aspects of the production process. You will create the critical tools and procedures to execute, capture, process, and present the data resulting from the end to end assembly and validation of our hardware across multiple vendors and sites. This role is hands-on and high-ownership. You will work closely across teams both internal and external to define the standards that will be used across our products to ensure the velocity and quality of our 1P hardware. You will own the implementation, deployment, and output of these systems as well their continued maintenance and SLAs. Location: San Francisco, CA (Hybrid: 3 days/week onsite). Relocation assistance available. In this role, you will: Design, develop, and maintain the software infrastructure for manufacturing process execution and data export. Own integration across internal customers and vendor systems and processes. Build and maintain the CI, release, and delivery pipeline of tooling to external partners. Build and maintain internal systems to ingest, process, deliver, and visualize critical data for internal teams and systems. Build system health monitoring, telemetry, remote d

PythonAWSLinuxRest
🔔

Get new infrastructure team manager jobs in United States by email

Daily job updates · Unsubscribe anytime