About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
Jobs in United States
Infrastructure Team Manager in United States
1,475 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
From $100K/yr
We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data structures,
About the Role We’re seeking an Associate General Counsel to lead commercial legal strategy and execution for OpenAI’s silicon initiatives and the semiconductor ecosystem that supports our AI infrastructure. This is a senior, highly cross-functional role for a lawyer who can advise business and technical leaders across the semiconductor development, manufacturing, and supply lifecycle. The role will partner closely with silicon engineering, infrastructure, strategic sourcing, supply chain, manufacturing, finance, partnerships, IP, policy, regulatory, and trade compliance teams to structure and negotiate strategic arrangements with semiconductor ecosystem partners, technology providers, and critical suppliers. This person will help establish the commercial frameworks needed to protect OpenAI’s technology and support resilient, scalable commercial arrangements. We’re looking for an experienced technology transactions lawyer with meaningful semiconductor industry experience who can translate highly technical and operational issues into practical commercial structures. The ideal candidate combines strong judgment, commercial creativity, and the ability to lead complex, high-value transactions in a fast-moving and complex global ecosystem. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. This role will: Lead commercial legal strategy and risk management for OpenAI’s silicon and related technology initiatives. Advise senior business, engineering, sourcing, and operational stakeholders on strategic relationships, commercial priorities, and complex transactions. Draft, negotiate, and advise on sophisticated development, manufacturing, supply, licensing, services, and strategic partnership agreements. Structure commercial arrangements that support long-term business and operational requirements. Advise on the commercial and operational issues that arise across the developmen
From $224K/yr
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Solutions Engineer (Enterprise Pre-Sales) Secure Every Identity, from Human to AI Agent About the Role At Okta, we believe that identity is the foundation of security and digital transformation. As a Senior Solutions Engineer, you will be the trusted technical advisor to our largest, most complex enterprise customers and a critical driver of our sales organization. We are looking for a highly strategic Pre-Sales professional whose consultative expertise, emotional intelligence, and enterprise sales acumen are their defining strengths. While technical agility is required, your primary focus will be owning the technical sales cycle by anchoring technical features to positive business outcomes. You will partner closely with Enterprise Account Executives to uncover top-of-mind business challenges—specifically around mitigating risk, reducing costs, and driving operational efficiency. By establishing value-based conversations, you will prove how Okta’s independent, neutral, and end-to-end identity platform can transform their architecture and secure the "Tech Win" on large-scale deals. What You'll Be Doing Master the Discovery Process: Leverage exceptional active listening to dig deep into customer pain points. You will uncover the 'why' behind the initiative, focusing on how Okta's vast pre-built integrations and scalable architecture can solve their most complex, diverse environmental challenges. Navigate Complex Organizations: Translate highly technica
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Director, Platform Engineering Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. With connections across more than 210 countries and territories, we are building a sustainable world that unlocks priceless possibilities for all. Role Summary Director, Platform Engineering At Mastercard, the pace of change across technology, AI, and the nature of work requires that we constantly push the boundaries for the technologies in building and operating infrastructure and applications globally. The Director, Platform Engineering is a critical technology and engineering leadership role within Operations Automation Program accountable for advancing our infrastructure and platform automation strategy through innovation and effective problem-solving. This role focuses on analyzing, coding, and delivering software an
From $145.7K/yr
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Platforms TPM organization partners with Experimentation, ML, LLM/GenAI, and Data Infrastructure teams to shape how Pinterest measures everything that matters — from product experiments to AI systems to the platforms everyone builds on. What you'll do: As a Staff Technical Program Manager for Measurement, you'll have the rare opportunity to shape how a company the size of Pinterest measures itself — turning a portfolio spanning experimentation, machine learning, generative AI, and data infrastructure into one coherent, high-impact program. Drive Pinterest's experimentation roadmap — accelerating how confidently and quickly teams can test, learn, and ship new ideas at scale. Own the program driving cost and compute efficiency across our ML systems, and help scale data science workflows into production-grade tooling. Lead cost optimization and
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: We are seeking an experienced Product Marketing Manager with a strong background in engaging developer audiences and delivering impactful go-to-market programs for native AI and enterprise companies. This role requires someone who is both technically savvy and strategic, with a proven track record of crafting compelling product narratives and building marketing assets that resonate with technical decision-makers. This role is specifically focused on our Model API product offering at Baseten. If you’re passionate about AI infrastructure, developer engagement, and simplifying complex technologies for real-world adoption, we want to hear from you. RESPONSIBILITIES: Positioning & Messaging: Develop clear and differentiated messaging that articulates the value of Baseten’s inference platform to developers and enterprise customers. Narrative Development: Shape how the market thinks about closed-to-open weights models and what matters most when building inference. Go-to-Market Strategy: Own the launch process for new features and products, collaborating closely with product, engineering, sales, and growth teams. Content Development: Create high-quality marketing assets, including white papers, technical blogs, demos, and customer case studies. Sales Enablement: Build resources and programs that empower our sales teams to effectively communicate Baseten’s capabilities and benefits. Market Insights: Understand the
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. ROLE We are seeking an experienced Product Marketing Manager with a strong background in engaging developer audiences and delivering impactful go-to-market programs for native AI and enterprise companies. This role requires someone who is both technically savvy and strategic, with a proven track record of crafting compelling product narratives and building marketing assets that resonate with technical decision-makers. This role is specifically focused on our post-training offerings at Baseten. We’re looking for someone who has experience in and is ready to learn more about all things post-training. If you’re passionate about AI infrastructure, post-training, developer engagement, and simplifying complex technologies for real-world adoption, we want to hear from you. RESPONSIBILITIES Positioning & Messaging: Develop clear and differentiated messaging that articulates the value of Baseten’s inference platform to developers and enterprise customers. Narrative Development: Shape how the market thinks about closed-to-open weights models and post-training ROI. Go-to-Market Strategy: Own the launch process for new features and products, collaborating closely with product, engineering, sales, and growth teams. Content Development: Create high-quality marketing assets, including white papers, technical blogs, demos, and customer case studies. Sales Enablement: Build resources and programs that empower our sales teams to eff
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. Posture Management & Visibility : Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build aut
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a visionary Principal Software Engineer to join our Compute organization and provide technical leadership for our Kubernetes infrastructure. You will drive the evolution of a platform that powers our global scale operations, transforming Kubernetes into a secure, reliable, and invisible foundation for our developers across our on-prem and public cloud fleet. Your mission is to balance cutting-edge innovation with rigorous platform stability, ensuring that internal and external customers have a seamless, high-performance experience at massive scale. You Will: Architectural Leadership: Serve as a technical lead for our Kubernetes ecosystem, setting the long-term architectural strategy for a platform that manages thousands of nodes and supports millions of concurrent requests. Customer Focus: Think deeply about how our internal customers consume and interact with compute, designing intuitive interfaces and tooling that simplify consumption of complex infrastructure services. Deep-Dive Engineering: Leverage deep expertise in Kubernetes internals, including custom controllers, operators, API server architecture, and etcd, to solve complex scaling bottlenecks and optimize our contr
From $100K/yr
We’re looking for Software Engineering Interns to help build and scale the systems that power Datadog’s observability and security platform. Interns contribute directly to real-world engineering challenges across backend, frontend, infrastructure, data engineering, and developer tooling while working alongside experienced engineers and mentors. You’ll help design, build, and improve systems that process and analyze massive volumes of metrics, logs, and application data in real time. Whether you’re interested in distributed systems, Kubernetes, AI-powered products like Bits AI, or developer platform tooling, you’ll work on meaningful projects that deliver impact to customers at global scale. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Contribute to production systems that process and analyze large-scale observability and application data in real time Build and improve distributed systems across backend infrastructure, developer platforms, and cloud-native services Help identify and solve performance, reliability, and scalability challenges in critical services supporting Datadog’s growing customer base Own and deliver technical projects from design through deployment with support from experienced engineers and mentors Develop technical expertise through hands-on experience with technologies such as Kubernetes, distributed systems, and cloud-native infrastructure Collaborate with fellow interns, mentors, and engineers while building software that delivers impact at global scale Who You Are: Pursuing a degree in Computer Science, Software Engineering, or a related technical field, or have equivalent practical experience Targeting a 2028 full-time start date Demonstrate strong computer science fundamentals, including data struc
Datadog is looking for a Senior Product Manager to help lead the evolution of our fleet and lifecycle management capability, the product surface that gives customers visibility into, and control over, the observability software running across their infrastructure. This capability manages the deployment lifecycle for core observability agents and OpenTelemetry collectors running on customer hosts and containers. The Senior PM will expand the scope of fleet capability to additional Datadog software components, making it the single place customers go to see everything running in their environment, at any version, in any deployment model, and to manage it remotely and safely at scale, for both human operators and, increasingly, AI agents acting on their behalf. This is a high-visibility, cross-functional role. You'll partner with multiple engineering teams and be responsible for defining and delivering a coherent, unified fleet experience across UI, API, and MCP for customers. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Own and evolve the product vision and roadmap for a unified fleet and lifecycle management capability spanning multiple product lines and deployment models. Define what "managed" means for each new software component as it's brought into fleet, balancing consistency of experience with the realities of each component's operational model. Drive a phased expansion plan, sequencing new components into fleet based on customer value, technical complexity, and dependency readiness. Partner closely with engineering leads across several teams to align on shared architecture principles to support disparate software components. Represent the voice of the customer for a capability that must work equally well for human operators using a UI and for AI
We’re looking for a Software Engineer to architect and build backend systems that enforce data privacy and automate compliance at scale. You’ll work closely with product, infrastructure, security, and legal teams to embed privacy-by-design into our data and access layers. This is a hands-on, high-impact role for an experienced engineer who is passionate about protecting user data while enabling innovation. What You’ll Do Design, build, and operate backend services that enforce policy-driven data access, lifecycle controls, and privacy protections. Develop distributed authorization and identity-aware enforcement mechanisms integrated directly into data services and control planes. Implement auditability, policy hooks, and enforcement observability to ensure compliance is continuously verifiable. Partner with Security, Legal, and Compliance to convert privacy requirements into scalable technical designs and developer-friendly APIs. Harden data platforms and backend services through schema-level controls and data handling constraints by default. Collaborate with infrastructure teams to ensure consistent enforcement across systems while minimizing duplicated implementations. Contribute patterns, libraries, and education that elevate trustworthy data access patterns across the organization. You Might Thrive in This Role If You Have 5+ years of industry experience building and operating backend or infrastructure systems in production. Strong software engineering fundamentals , with fluency in at least one major programming language (e.g., Python, Go, Rust, C++, Java). Experience with distributed authorization, RBAC/ACL systems, encryption-based access, or policy engines. Familiarity with global privacy regulations and their architectural implications. Ability to influence and collaborate with teams across legal, compliance, product, and engineering. A bias toward practical, impactful solutions that balance privacy protections with product needs. Nice to Have Experience wi
About the Role We’re looking for a senior Strategic Finance Lead to help shape and execute OpenAI’s financial strategy across some of the company’s most important long-term investments and growth priorities. This high-impact role sits at the intersection of corporate finance, capital markets, treasury, and cross-functional execution. In this role, you will: Shape financial strategy for major long-term investments, capital structure decisions, and broader strategic decision-making. Build rigorous financial models and scalable frameworks to evaluate strategic initiatives, funding options, tradeoffs, and risk. Structure and help execute complex strategic initiatives in partnership with cross-functional teams. Translate ambiguous technical and business inputs into clear recommendations and decision-ready materials for executives, the board, and other senior stakeholders. Lead high-priority cross-functional workstreams from concept through execution, bringing strong judgment, ownership, and communication in fast-moving situations. Help build an AI-native finance organization by applying AI to improve workflows, decision-making, and execution across finance processes. You might thrive in this role if you have: 12+ years of experience across strategic finance, corporate finance, capital markets, investment banking, infrastructure finance, or related fields. A strong understanding of capital structure, financing strategy, debt markets, and large-scale investment evaluation. Experience operating in high-growth, high-complexity environments with the ability to navigate ambiguity and move quickly. Exceptional financial modeling, strategic thinking, and problem-solving capabilities. Strong executive presence with the ability to communicate clearly across investors, executives, and technical stakeholders. A high-agency operating style that combines strategic perspective with a willingness to build directly. Excitement around using AI as a core operating advantage, with a mindset
Other cities to consider
More places hiring for this role
Get new infrastructure team manager jobs in United States by email
Daily job updates · Unsubscribe anytime