We are looking for a Program Manager to join our People Team, reporting to the Chief of Staff to the Chief People Officer. In this role, you will drive large-scale, high-impact, and cross-functional programs that shape and evolve our People products and infrastructure. You will operate at both a strategic and executional level, partnering across the People Team, and with business and functional stakeholders across Finance, Legal, IT, and business leadership to deliver transformative initiatives that support Datadog’s growth. This is a highly visible role requiring strong program management, systems thinking, and the ability to connect initiatives across the organization while improving operational effectiveness and employee experience. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead end-to-end delivery of complex, cross-functional People programs, from ideation through execution and iteration Partner with senior stakeholders across the business to define program goals, success metrics, and roadmaps Deliver People products through scalable processes and frameworks that improve employee experience and organizational effectiveness Identify dependencies, risks, and tradeoffs across initiatives, proactively driving alignment and decision-making Translate strategic priorities into actionable plans, ensuring clear communication and accountability across teams Establish program governance, reporting, and operating rhythms to track progress and outcomes Synthesize insights across multiple workstreams to inform executive-level updates and recommendations Continuously improve program management practices within the People Team Who You Are: 5+ years of experience in program management, operations, consulting, or a related field, preferably
Jobiba hiring network
Infrastructure Team Manager Jobs
4,815 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. We are seeking a strategically-minded, technology-focused, and customer-centric Engineering Manager to lead one of our Infrastructure teams here. You will lead a team responsible for building, operating, and scaling the cloud infrastructure and platform systems that underpin CLEAR’s services, ensuring reliability, performance, and security across our environments. A successful candidate brings strong experience in cloud infrastructure, distributed systems, and operational excellence, along with a solid foundation in software engineering. You are an effective communicator who can lead complex infrastructure initiatives from inception through delivery, and thrive in fast-paced environments. This role requires a focus on building resilient, scalable systems, driving automation, and leading and developing high-performing engineering teams. What you'll do: Hire, develop, and grow engineering talent through coaching, mentorship, performance management, and career development planning Set clear goals and expectations, provide regular feedback, and foster accountability across the team Own and execute the roadmap for cloud infrastructure and platform engineering, and reliability initiatives Design, build, and operate a scalable, secure, and highly available cloud platform infrastructure Drive automation across infrastructure provisioning, deployment, and operations to improve efficiency and reduce manual overhead Establish and enforce best practices for system reliability, observability, incident response, and disaster recovery Partner with eng
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE At Baseten, we’re looking for a Technical Program Manager to drive our most complex, cross-cutting infrastructure programs. This role will operate across all domains of AI infrastructure, from the GPUs up to the multi-cluster orchestration layer. This is an execution-first role. The work is less about owning a single system and more about imposing order on ambiguity: standing up the right structures, driving decisions to closure, and making sure nothing falls through the cracks across dozens of stakeholders. If you take satisfaction in turning a chaotic, half-defined initiative into a predictable, well-governed program, this role is for you. RESPONSIBILITIES Own complex migrations end to end. Lead large-scale infrastructure migrations across teams and domains. This will involve scoping the work, sequencing dependencies, managing risk, and driving them to completion without surprises. Drive process across infrastructure. Establish and run the operating rhythms that keep programs healthy: planning cadences, status reporting, decision logs, risk reviews, and escalation paths. Make the process light enough that teams adopt it and rigorous enough that it actually works. Help managers build the right structures. Partner with engineering managers and leads to design the team structures, ownership boundaries, and working models a program needs to succeed. Spot gaps in accountability before they become problems. Own fo
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! About the Role We're looking for an Engineering Manager to lead our Deployment Engineering team in EMEA. This isn't a typical management role — we need someone who leads from the front, gets their hands dirty, and drives impact. You'll manage a team of Forward Deployed Engineers who are on the front lines of deploying Cohere's North platform into customer environments. You should be ready to be a force to be reckoned with. Location: UK (can be remote but based in UK) What You'll Do Lead and mentor a team of Forward Deployed Engineers across EMEA Drive end-to-end deployment of North in private cloud and on-premises environments Take ownership of customer success from technical implementation through delivery Collaborate closely with Product, Engineering, and Sales to shape how we deliver AI to enterprises Mentor your team on cloud infrastructure, Kubernetes, and enterprise-grade deployments Optimize performance for OpenSearch, databases, and other K8s services Define scaling guidelines for GPU and CPU compute resources Build processes & technology that scales — we're growing fast. What We're Looking For 5+ years of experience
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! We're looking for an Engineering Manager to lead our Deployment Engineering team. This isn't a typical management role — we need someone who leads from the front, gets their hands dirty, and drives impact. You'll manage a team of Forward Deployed Engineers who are on the front lines of deploying Cohere's North platform into customer environments. You should be ready to be a force to be reckoned with. Location: North America (remote-first) What You'll Do Lead and mentor a team of Forward Deployed Engineers Drive end-to-end deployment of North in private cloud and on-premises environments Take ownership of customer success from technical implementation through delivery Collaborate closely with Product, Engineering, and Sales to shape how we deliver AI to enterprises Mentor your team on cloud infrastructure, Kubernetes, and enterprise-grade deployments Optimize performance for OpenSearch, databases, and other K8s services Define scaling guidelines for GPU and CPU compute resources Build processes & technology that scales — we're growing fast. What We're Looking For 5+ years of experience in software engineering with demonstrate
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this team? The GPU Clusters team builds and operates the superclusters that train Cohere’s frontier models. We sit at the intersection of hardware, distributed systems, and AI research. We work with cloud providers, researchers, and other infrastructure teams on problems few companies get to take on. As an Engineering Manager, you’ll lead a team of engineers who care deeply about GPU infrastructure. You’ll set technical direction, grow people, and help the company scale a rapidly growing compute footprint. As an Engineering Manager, you will: Hire, mentor, and grow a team of GPU infrastructure engineers , including performance, career development, and technical guidance on hard infrastructure problems Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance Partner with researchers and ML engineers so the training and inference stack works well on new GPU architectures Work with cross-functional stakeholders such as Capacity, Finance, Legal, Security, and other infrastructure teams on planning, cost, compliance, an
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As Manager / Senior Manager, Cell Infrastructure , you'll help shape the foundation that lets GitLab run reliably across multiple cloud providers in the agentic era. As more enterprise customers adopt AI-driven development workflows, GitLab needs a cell-based infrastructure that can scale horizontally, support strong reliability, and operate with clear cost efficiency. In this role, you'll report to the VP Engineering, Platform Scale & Architecture and lead a 0 to 1 effort to build the control layer that determines how GitLab cells are provisioned, placed, and operated across clouds. This is a high-im
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Observability team builds the infrastructure that empowers engineers to understand, operate, and improve the Roblox platform and ecosystem. Our team owns the end-to-end observability stack across telemetry, distributed tracing, logging, profiling, storage systems, and developer-facing visualization tools. We are looking for an Engineering Manager to lead the next generation of AI-powered observability platforms. In this role, you will help build intelligent systems that leverage AI to revolutionize CI/CD, testing, and DevOps workflows — enabling engineers to move faster, improve reliability, and operate large-scale distributed systems with greater efficiency and confidence. This is a highly impactful leadership role at the center of Roblox infrastructure. Your work will directly improve developer productivity, platform reliability, and operational excellence across the company. You will partner closely with infrastructure, product engineering, and AI platform teams to shape the future of developer tooling and autonomous operations at scale. You Have 3+ years of engineering management experience with a proven track record of hiring, mentoring, and growing high-performing teams. Strong ex
We’re looking for an experienced sourcing leader to join the Datadog Procurement team and help grow the Strategic Sourcing group. Make an impact by partnering closely with senior technology and business leaders, owning our fastest-growing AI, neocloud, and inference spend end-to-end, and continuing to prove the value that Strategic Sourcing brings to the organization. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the AI spend category end-to-end, covering foundation-model and AI APIs, AI development and productivity tools, and AI-enabled SaaS (with neocloud and inference compute as an emerging area), while developing and executing category strategies aligned to business objectives Champion AI and automation adoption across the sourcing function by identifying tools (e.g. Claude, ChatGPT) and building repeatable workflows that make sourcing faster, smarter, and more scalable Lead high-value AI vendor negotiations spanning foundation-model and API agreements, AI development and productivity tools, and AI SaaS, structuring seat- and consumption-based pricing across net-new purchases and strategic renewals, and building neocloud and GPU-capacity capability as the category grows Partner with Engineering and Finance to turn architectural and consumption trade-offs into commercial business cases, forecasts, and savings targets Build pricing and consumption models and apply FinOps discipline to quantify buying scenarios and uncover savings across a fast-moving spend base Work alongside executive and senior technology leadership as a trusted advisor on AI and neocloud investment decisions, influencing strategy and commercial trade-offs at the leadership level Track category KPIs (savings, pipeline, cycle time), monitor AI market, vendor, and pricing trends,
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta Identity Governance (OIG) organization is looking for a Software Engineer Manager to join our team — OIG is Okta’s Identity Governance and Administration solution that is directly responsible for one of the most critical and visible workflows in enterprise identity: how employees request, approve, and gain access to the resources they need. Opportunity As a Software Engineer Manager on the OIG team, you will manage the team that is responsible for designing and building Identity Governance product features — spanning the across different access governance personas. You will lead a team of elite engineers, grow the team, and drive impact and customer satisfaction. You will drive engineering best practices, productivity enhancement, and make our elite engineering team shine. You will also work collaboratively across engineering, product, and design to deliver features that are secure, scalable, and delightful to use. This is a rare opportu
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies, from the world's largest enterprises to the most ambitious startups, use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe's Infrastructure organization builds and operates the shared platforms that make it possible for engineering teams to build, deploy, run, and scale reliable products. The organization owns foundational capabilities across service networking and discovery, safe feature and configuration changes, data storage and orchestration, and batch computing. Our platforms support a global engineering audience and power systems at significant scale. This includes the control plane for secure service-to-service communication across tens of thousands of microservices, service and feature deployment infrastructure that helps teams introduce and recover from changes safely, infrastructure for running Stripe's containerized services, and data platforms for storage, pipeline orchestration, and large-scale processing. Our batch compute platforms support technologies including Apache Airflow, Apache Spark, Apache Iceberg, Apache Hadoop, and Apache Celeborn. As an Engineering Manager, you will lead or strongly influence teams working on business-critical infrastructure with broad internal adoption. You will balance reliability, scalability, security, developer experience, and operational rigor while helping Stripe move quickly with confidence. You will partner closely with product, security, data, support, and adjacent infrastructure leaders to set direction, resolve cross-team dependencies, and turn recurring needs into durable platform capabilities. Wh
Anyscale Platform Engineering Leader About Anyscale: At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI , Uber , Spotify , Instacart , Cruise , and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role: Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless. In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges. You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components. You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful. We'd love to hear from you if you have: Solid engineering management experience leading produ
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? As Cohere continues to grow, the size and complexity of our programs has continued increasing over time. To help us manage this, we are looking to bring in exceptional Technical Program Managers (TPM) to manage this! At Cohere, TPMs are pivotal to our success, being seen as operational experts specializing in cross-functional efforts. Great TPMs at Cohere are well-rounded, executing with precision at the tactical level, while partnering with senior leaders and shaping the big picture at a strategic level, and the opportunities are endless in growing the scale and impact of their work. Cohere’s TPMs are often described by their stakeholders as organized, execution-oriented, pragmatic, and adaptable to the needs of the company and the programs they lead. As a reward for the depth and breadth of expertize that you bring to the table, you will get to work alongside some of the most talented engineers in the world on building cutting edge AI technology. No day will be the same as the one before, and the dynamic high growth environment will be absolutely perfect for anyone excited about solving novel problems, bringing
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver effective performance ads to our users, and more business values to our advertisers. We’re looking for an EM to lead a team of exceptional ML infrastructure engineers, build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You Will: Lead strategic planning and roadmap execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference. Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature). Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment. Recruit, mentor, and grow a high-performing team of ML infrastructure engineers. You Have: 5+ years of experienc
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Infrastructure Platform and Shared Services Team Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling. As the Sr. Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, Observability, automation platform & tooling. What you’ll be doing Lead the Infra platform and shared services org and various initiatives across SRE & Infrastructure organization. Build a world-class observability platform and monitoring capabilities enabled with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and intuitive self-service capabilities. Own the design and operation of scalable, self-service Cloud infrastructure platforms (e.g. Observability Platform, SRE Productivity, deployments, and Edge Infrastructure) Lead, mentor, and grow a high-performing team of engineers and managers across SRE and infrastructure shared services domains. Perform engineering design evaluations and ensure the completion of projects within resource,
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime