Jobs in United States

Infrastructure Team Manager in United States

1,531 active opportunities · Updated October 2026

Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization’s mission is to become the leader in civil immersive online communities. We systematically detect, remove, and prevent problematic accounts, content and behavior, and we make Roblox accounts secure and free from compromise. We cover a broad area of the tech spectrum, including machine learning, classifiers for 3D models, experimentation, automation, detection workflows, and AI-powered text filters. Aligned and partnering with product teams, we use this tool-belt to discover new opportunities, influence and shape the product roadmap and prioritization, build safety products, and measure the impact on our community of users and developers. In doing so, we keep Roblox safe, civil, and inclusive, and we foster positive relationships between people around the world. WHY Safety Data Infrastructure To pro-actively find bad actors and protect good users, Roblox needs to ingest enormous amounts of data and make it usable both to our automated detection systems and for human moderators and customer support agents doing investigations. The Safety Data Infrastructure team is addressing these needs w

PythonAWSGitMicroservices
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale. In this role, you will: Design, build, and operate the distributed storage system that feeds model training and evaluation. Run this system multiple on Kubernetes clusters at petabyte scale. Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success You may be a good fit if you have: Strong storage fundamentals,

PythonKubernetesGitRest
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.6%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry provides tools that help developers find and fix issues in their applications. The Developer Infrastructure team owns the systems that make every engineer at Sentry more effective at shipping high quality software: developer environments, CI/CD for our open source and closed source codebases, the golden path for building new services, our SDK and library publishing tooling, and the metrics and dashboards that keep it all healthy. Our goal is straightforward: developers should spend their time thinking about the software they build for customers, not the software they use to build it. That means everything from local and cloud development environments, to the CI/CD pipeline that gets code out safely and quickly, to the tooling that catches flaky tests before they cost someone a day. As AI coding agents become part of how engineers work, we're also investing in making sure our environments, CI, and deployment systems support that shift. As a Software Engineer on Dev Infra, you'll help build and scale this infrastructure end to end. In this role, you will Build and maintain the tooling that powers local and cloud-based developer environments Improve CI for our codebases and CD for our deployment pipeline, keeping both fast and reliable as the org and its infrastructure grow Define and evolve the golden path for building new services and libraries at Sentry Contribute to our SDK and library publishing tooling and release processes Build the metrics, dashboards, and flaky test detection that give the org visibility into deployment health and CI reliability Build tooling that gives engineers, and the AI codin

PythonDockerCI/CDRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About OpenAI OpenAI is dedicated to ensuring that artificial general intelligence (AGI) benefits all of humanity. Our mission requires building not only world-class AI models, but also the infrastructure that enables those models to be deployed reliably, efficiently, and at global scale. As demand for AI continues to grow, we are expanding the ways OpenAI can bring high-performance inference capacity online across a diverse hardware ecosystem. About the Team The GPT Infrastructure team builds software that turns advanced inference and optimization research into production products. One focus is enabling strategic infrastructure partners and accelerator vendors to qualify and onboard new compute without a bespoke porting and optimization effort for every hardware platform. We build the control planes, APIs, secure partner-side execution environments, evaluation systems, artifact pipelines, and operational tooling that make these workflows repeatable and trustworthy. The work sits at the intersection of distributed systems, AI inference, compilers and runtimes, performance engineering, security, and external partnerships. About the Role We are seeking an experienced systems generalist who can work comfortably across the stack to help build an automated inference optimization platform. Given a workload, target hardware profile, compiler and runtime context, and a trusted verifier, the system runs durable optimization campaigns that generate, compile, execute, grade, and improve candidate kernels, runtime configurations, and serving-stack changes. You will design both the OpenAI-hosted control plane and the partner-side software that evaluates candidates on real accelerator hardware. The product must keep long-running workflows reliable, make performance results reproducible, and maintain clear trust boundaries around sensitive model and hardware information. This is a deeply cross-stack role, combining strong software engineering fundamentals with systems thinking and

PythonAWSLinuxRest
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $287.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on the Privacy Infrastructure team, you’ll design and build large-scale systems that operate across Roblox’s platform and data ecosystem. You’ll tackle complex distributed systems challenges, develop reliable and scalable infrastructure, and work across engineering teams to deliver foundational capabilities that serve millions of users. This is an opportunity to take on high-impact technical problems at Roblox scale while helping shape the next generation of our infrastructure. You Will Design, build, and scale reliable infrastructure and platform solutions that protect user data and support the needs of a global platform. Develop foundational capabilities for emerging technologies, including agentic AI workflows, with a focus on safe and responsible data access. Partner with engineering teams across Roblox to integrate scalable data protection capabilities into their systems and development workflows Drive technical strategy and architecture for complex, cross-cutting challenges spanning data, infrastructure, privacy, and security. You Have 7+ years of proven experience as a software engineer. Expertise in Python or Golang; Strong understanding of system

PythonAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You’ll own the hands-on and automation work that brings WAN, fiber, carrier, and cloud-interconnect circuits into service. Partner with network engineers, fiber providers, cloud service providers, colocation teams, and data-center technicians to move each connection from ordered and patched to verified, stable, and ready for handoff. You’ll own Layer 1 troubleshooting and circuit bring-up while building workflows that translate reliable system or model output into precise, approved technician actions, capture field feedback, and drive each connection to a green-port handoff. The right person combines strong physical-networking judgment with practical automation skills: patch-panel and port mappings, optics and light levels, provider coordination, structured operational data, API or scripting workflows, and human-in-the-loop LLM tooling. Responsibilities Own Layer 1 activation and restoration for carrier circuits, dark fiber, wavelengths, Ethernet handoffs, and dedicated cloud interconnects across data centers and points of presence. Reconcile complete A-side/Z-side as-builts: circuit IDs, LOAs/CFAs, carrier demarcations, MMR/ODF/MDF and patch-panel positions, fiber pairs, cross-connects, optics, and device ports. Investigate no-light, low-light, wrong-port, link-flap, and error-rate issues across providers and CSPs; isolate continuity, dirty connectors, polarity, incorrect patching

AWSAzureRestAI
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA is seeking an a PCB Library Engineer to join our PCB Design Infrastructure team. In this role, you will help develop and maintain the PCB library assets used across NVIDIA's Data Center, AI, Networking, Automotive, and Graphics products. Working alongside experienced PCB designers, library engineers, mechanical engineers, manufacturing engineers, and component engineers, you will create and validate component footprints, schematic symbols, mechanical components, panel definitions, and other critical design assets that enable successful product development. This position provides an excellent opportunity to build expertise in PCB design, manufacturing, component engineering, and design automation while supporting some of the most advanced computing platforms in the world. What you'll be doing: Develop PCB footprints, padstacks, schematic symbols, and mechanical library content using Cadence PCB design tools. Review component datasheets, package drawings, and engineering specifications to create accurate design libraries. Support library verification, release, and documentation processes. Partner with PCB design, mechanical engineering, manufacturing engineering, and operations teams to resolve library-related issues. Learn and apply industry standards, including IPC requirements, DFM, DFA, and DFT principles. Support quality initiatives to ensure library content is accurate, manufacturable, and scalable. Participate in continuous improvement and automation efforts within the library environment. Develop a strong understanding of PCB fabrication, assembly, and component technologies. What We Need to See: BS degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Manufacturing Engineering, or a related field or equivalent experience. <p

S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Snowflake is about empowering enterprises to achieve their full potential, and people too. With a culture that’s all in on impact, innovation, and collaboration, Snowflake is the sweet spot for building big, moving fast, and taking technology, and careers, to the next level. About the Role The Cortex Code team is building the future of coding agents for working with data. See our flagship product in action: Cortex Code in Action: Live Demos + AMA . Your work will directly impact how developers and businesses build with data. You'll own the full AI engineering lifecycle: design, prompt/tool engineering, evals, deployment, measurement, and optimization. You'll work with a small, high-powered modeling and infrastructure team. What you will do in this role: Own features end-to-end for Snowflake Cortex Code products. Build agentic workflows, coding harnesses, evaluation pipelines. Build enterprise-grade context engineering: function calling, tool schemas, guardrails, agent teams, and verification/repair. Partner with product and infra: translate customer problems into products and experiments. Collaborate with infrastructure teams to productionize improvements. Work with an elite team of engineers towards building great products Requirements: Bachelor’s degree in Computer Scienc

TypeScriptPythonAIGo
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -83%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a member of the Capacity Strategy & Operations team, you will sit at the intersection of supply intelligence, demand forecasting, and cross-functional execution, turning a complex, fast-moving hardware market into a predictable, reliable foundation for our customers and internal engineering teams. This is not a purely analytical role. You will own the end-to-end capacity planning process: from translating customer commitments and growth forecasts into concrete supply requirements, to coordinating fulfillment across vendors, finance, and the infrastructure team, to building the systems that make all of this repeatable and scalable. When supply is constrained and tradeoffs are unavoidable, you are the person in the room who can model the options, make a clear recommendation, and drive alignment fast. You are a strong fit if you have operated at the intersection of strategy and execution before — someone who is equally comfortable building a capacity model in a spreadsheet and running a cross-functional war room when a customer deployment is at risk. EXAMPLE INITIATIVES Demand-Supply Alignment Framework: Build and own the process that translates customer pipeline, signed commitments, and growth projections into a forward-looking GPU demand signal — so the team is never caught flat-footed when a customer scales faster than expected. Constrained Allocation Playbook: Define the decision framework for how Basete

Machine LearningAIGoExcel
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.6%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the Role This isn’t a typical engineering role. You won’t be embedded in a single product team or siloed in one product area. Instead, you’ll sit within Platform Engineering, own the AI-assisted coding domain, and work across all of engineering at Sentry, focused specifically on how AI coding agents participate in our software development lifecycle. For AI coding agents to work well in our repo, the internal systems they depend on need to be accessible via API, not locked behind UIs that require human interaction. Right now, many of those systems aren’t agent-ready. You’ll audit and prioritize that gap, expose those systems programmatically, and build the connections that let tools like Claude Code operate on them end-to-end. From there, the scope expands to improving the quality of AI-generated pull requests and automating the engineering work that’s important but consistently deprioritized. You will look from context engineering standpoint to see what to send to our model; you will look from harness engineering standpoint to see the tools it can use, the permissions it has, the state it carries forward, the tests it has to pass, the logs you capture, the retries, checkpoints, guardrails, and evals. You’ll work closely with the dev infrastructure team as your home base, then collaborate across every product team coding in our repo once the tooling foundation is in place. It’s a broad role with real impact, and the work you do will directly change how Sentry engineers ship software. What You’ll Do Audit Sentry’s internal developer systems and make them API-ready for AI agents. You’ll prioritize and drive the work of ex

CI/CDMachine LearningAIGo
I
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -90.5%
Quick readStrong listing-quality and freshness signals

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Our backend systems power the clients used by millions of customers every year to buy their groceries online. These systems must also support tight integration with the largest retailers in the US and Canada. Engineering at Instacart provides the opportunity to work on challenging scaling problems while also designing the features that will define our industry. You will learn how to build in an open collaborative environment serving millions of requests daily. Finance data engineering is part of the Data Infrastructure team, working closely with accounting, billing & revenue teams to support the monthly/quarterly book close, retailer invoicing and internal/external financial reporting. The team plays a critical role in defining how financial data is modeled and standardized for uniform, reliable, timely and accurate reporting. This is a high impact, hi

PythonSQLAIGo
S
📍 Menlo Park, California, United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Cortex Code is Snowflake’s coding agent for building with data. It ships inside the platform that thousands of the world’s largest enterprises — including a large share of the Forbes Global 2000 — run their data on, which means the quality of this agent is felt by the data teams behind a meaningful slice of the global economy. We are taking coding agents from impressive demos to tools Data Science and Engineering teams depend on every day, and we hold them to a rigorous, public bar: see our data engineering agent benchmark . About the Role This is a measurement-first role that owns the quality and efficiency of Cortex Code end to end: how good the agent is, how much it costs to run, and how reliably it behaves in production. You will take agents from research capability to real, measurable user value — turning fuzzy “the agent feels worse” signals into hard metrics, running the experiments that move them, and shipping the changes that stick. You will work on a small, high-powered modeling and infrastructure team where your work reaches every developer building on Snowflake. What you will do in this role Take agents from prototype to production: design and refine agent behaviors for real coding and data-engineering workflows, and make them reliable enough to depend on. Own a

TypeScriptPythonAIGo
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -87.5%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role: Join our Infrastructure Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Staff Infrastructure Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability. We are seeking Staff Infrastructure Engineers who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. You Will: Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios. Optimize Performance and Infrastructure: Collaborate with core infrastructure and product teams to performance tune and optimize our cloud deployments (Kubernetes, Docker, GCP). Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions. Elevate Developer Experience: Design and implement improvements to our build, test, and deployment systems to make software delivery faster, safer,

PythonGCPDockerKubernetes
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Senior Software Engineer to own key components of our AI native External Observability Platform . In this role, you will contribute to the technical road map for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex technical projects, and skilled at collaborating with the brightest technical minds in the industry. Key Responsibilities Develop and Scale Distributed Infrastructure: Design and implement key components of Snowf

JavaVueAWSAzure
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams. Our work spans the entire stack: capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience that lets teams use enormous compute systems with confidence. At this scale, small improvements to communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. We are hiring across Compute Infrastructure rather than for a single narrow team, and we use this opening to match strong engineers to the problems where they can have the most leverage. About the Role We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you. Depending on your background and interests, you might work close to hardware, close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You could help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, t

AWSKubernetesRestAI
🔔

Get new infrastructure team manager jobs in United States by email

Daily job updates · Unsubscribe anytime