A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Are you passionate about engineering quality, performance, and increasing the impact of engineers around you? Software Engineers at Palantir build software at scale to transform how organizations around the world use data. As an intern within Palantir’s Foundations organization, you’ll have the opportunity to grow more quickly than you ever imagined, as you build the shared infrastructure that underpins the Palantir Foundry, Palantir Gotham, and Palantir Apollo platforms, and drive investments to improve the velocity and quality of our engineering. Teams within Palantir’s Foundations organization are made up of a small number of engineers, each focused on one of four major categories of our infrastructure: • Backend Infrastructure: Maximizes the productivity of our backend developers and ensures Palantir’s platforms have performant and consistent RESTful services. Think: making the “easy way” the “right way” when developing backend services, including designing infrastructure to build hundreds of micro-service repos performantly, or to ensure we keep reliable audit logs of everything users do in our platforms. • Developer Infrastructure: Operates the systems and services that underpin all aspects of our developer ecosystem, including off-the-shelf tooling like GitHub and custom tooling for managing automated changes across hundreds of repositories. • Frontend Infrastructure: Maximizes frontend developer productivity across the entire frontend development stack, from the developer experience in the IDE to the final user experience in the browser. Think: the core infrastructure required to develop and serve our frontends (including features flags, inter
Jobiba hiring network
Infrastructure Team Manager Jobs
4,730 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions for a platform serving 200M+ daily active users. As a Principal Software Engineer in our Data Infra org, you will be the primary technical leader driving the strategic vision, long-term architecture, and massive scalability of our distributed data platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role operates under high ambiguity, demanding unparalleled ownership to redefine the limits of infrastructure handling exabyte-scale workloads, and providing a unique opportunity to lead the future evolution of our global data ecosystem. You Will: Define Multi-Year Technical Strategy: Own and drive the end-to-end architectural vision for Roblox's core data platforms spanning Kafka, Flink, Spark, Trino, Druid, Airflow, and Data Catalog systems. Turn multi-year company strategies into concrete, production-grade infrastructure blueprints. Lead Cross-Functional Alignment: Partner closely with executive leadership, platform governance, data science, and product e
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Software Engineer for Infrastructure Security you will be a part of the Information Security organization and report to the Senior Manager of Infrastructure Security. You will help shape the future of Platform Security at Roblox. We work closely with Production IAM, Network Security, and Cloud Security at Roblox. You Will: Identify security gaps and threats in our cloud and on premise infrastructure, partnering with Governance Risk and Compliance teams to create standards and policies along the way. This will help Roblox meet regulatory and compliance requirements. Harden our infrastructure by introducing secure by default configurations, designs and guardrails for all developers at Roblox. Own and drive solutions that enable Roblox engineers to design, build, and use infrastructure securely at scale. Work closely with other InfoSec teams (AppSec, D&R, GRC, CorpSec, CloudSec, NetSec) and partner with engineering teams across Roblox, specifically the Infrastructure organization, to ensure the secure outcomes of security and product driven initiatives. You Have: 5+ years of experience writing code and/or relevant technical experience. Experience with
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox's data infrastructure processes petabytes of data daily, powering analytics, ML, and product decisions. As a Senior Software Engineer in our Data Infra org, you will design, build, and scale the distributed data infrastructure platforms that power Roblox. You will own and drive the next-generation architecture of our core platforms, which span Kafka, Flink, Spark, Trino, Druid, Airflow and Data Catalog. This role combines high ambiguity and ownership to push the boundaries of what our infrastructure can handle at massive scale, giving you the unique opportunity to steer the evolution of the data landscape. You Will: Own and Scale Core Platform Components: Take responsibility for the design, architecture, and implementation of 1–2 key data platform frameworks within our stack Collaborate and Align: Partner with infra, data science, and product engineering teams to ensure your target platform's capabilities are directly guided by platform governance and product requirements. Optimize Performance at Scale: Dive deep into engine internals, query planning, state management, memory optimization, serialization efficiency to maximize throughput and reliability under heavy load. Drive Infrast
About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . This role focuses on advancing the science and systems behind ML measurement, feature understanding, and causal inference at scale. The work spans areas such as production feature importance platforms, observational causal estimation in Pytorch, large-scale proxy metric development, and data-driven approaches to ML infrastructure efficiency. We're looking for an enthusiastic individual contributor to perform high-impact technical work across this space. This person will drive foundational innovations, own the end-to-end design of production ML systems, establish rigorous methodological standards, and partner cross-functionally to turn successful research into durable platform capabilities that raise the ceiling for the entire ML organization. What you’ll do: We are looking for an experienced and highly capable Data & Applied Scientist
Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Working with teams across the organization to understand pain points in their infrastructure usage to find common ideas and work to create solutions which span multiple domains; Define and produce high quality written proposals, communications and documentation. Execute on technical programs that require deep systems and engineering level knowledge; Partner with Engineering Managers, Tech Leads, Engineers and other Technical Program Managers to define, scope and drive large programs to conclusion; Play a key part in shaping the technical design, predicting technical roadblocks by collaborating with engineers, and identifying trade-offs; Develop, implement, and iterate on program management techniques, frameworks, and KPIs to achieve goals with well defined success criteria; Elevate the execution muscle of engineering teams around you; Train them to be better at delivery where needed; Help influence peers / stakeholders and build consensus while dealing with ambiguity; Leverage data and acquired knowledge to drive strategic decisions at an engineering leadership level; Create widely circulated plans, driving consistency, clarity and building alignment across teams; Operationalize and execute critical cross functional programs spanning multiple engineering organizations for Infrastructure (Developer Infrastructure, Core Infrastructure, Service Infrastructure); Shape technical design, predicting tech
About the Role We are seeking a Cloud Infrastructure Engineer to help design and evolve the platforms that power OpenAI’s products. In this role, you will be a hands-on technical leader, driving the architecture, scalability, reliability, and security of critical infrastructure systems. You will help define how we build and operate infrastructure at the next order of magnitude, while influencing technical direction across teams. This role is both deeply technical and highly strategic, requiring strong ownership, sound judgment, and the ability to partner effectively across engineering, product, and research organizations. In this role, you will: Design and build scalable, reliable, and secure infrastructure platforms that power OpenAI products Evolve cloud infrastructure abstractions that enable rapid product development across teams Architect systems to support significant growth, performance, and operational complexity Improve server orchestration, networking, distributed systems reliability, and infrastructure security posture Influence technical direction and infrastructure strategy across multiple teams Partner closely with product, research, and engineering teams to align infrastructure with evolving needs Own operational excellence, including participation in on-call rotations, incident response, and production readiness Mentor engineers and raise the overall technical bar of the organization Contribute to a culture of high ownership, low ego, and thoughtful collaboration You might thrive in this role if you: 8+ years of experience building and operating large-scale infrastructure systems Deep expertise in Kubernetes and container orchestration at scale Strong experience designing cloud abstractions and platform infrastructure (AWS, GCP, Azure, or similar) Proven track record of leading complex technical initiatives across teams Experience operating highly reliable, secure, and scalable distributed systems Security engineering experience or security backgroun
About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Role We are looking for a Principal Product Manager to lead the strategic evolution of the Auth0 Platform Infrastructure. In this highly visible, high-impact leadership role, you will be the visionary force behind the core infrastructure, consumption models, and platform capabilities that power Auth0’s immense global scale. You will own the multi-year strategy for critical platform pillars, with a heavy emphasis on redefining our rate-limiting and platform consumption frameworks. As a Principal product leader, you will navigate multi-dimensional scaling challenges (including the paradigm shift of autonomous AI agents), influence executive stakeholders, and ensure our platform remains the most resilient, secure, and performant identity solution for the world’s largest enterprises. What You’ll Do Drive the Long-Term Platform & Consumption Strategy: Define and own the multi-year roadmap for Auth0’s core infrastructure. Align deeply technical platform architecture with broader business objectives, Go-To-Market strategies, and overarching revenue goals. Spearhead the Strategic Evolution of Rate Limits & Monetization: Transform rate limits from a protective mechanism into a core business strategy. Pioneer dynamic consumption models, proactive policy-aware observability, and intelligent throttling strategies that align platform costs with customer value, reduce escalations, and unlock new SKUs. Architect Global Cloud Resilience & Scale: Champion ov
About Graphcore Graphcore is a global leader in artificial intelligence computing systems. We design advanced semiconductors and data center hardware that deliver the specialized processing power needed to advance AI while improving the efficiency required for broad adoption. As part of SoftBank Group, Graphcore belongs to a family of companies developing some of the world's most transformative technologies. Our AI Engineering Campus in Austin plays an important role in building the future of AI computing. The Opportunity As Technical Services Director, you will lead the teams that operate and evolve Graphcore's engineering labs, high-performance computing (HPC) platforms, and data center environments globally. You will be accountable for reliable, secure, cost-effective infrastructure that supports demanding engineering, AI, silicon-development, and validation workloads. This role combines people leadership, infrastructure strategy, operational excellence, capacity and financial planning, procurement, and program delivery. You will partner with Engineering, Information Technology, Security, Finance, Facilities, Supply Chain, customers, and external suppliers. The position is based onsite in Austin and requires travel to company facilities, data centers, and supplier locations, including international travel. What You'll Do Lead, recruit, mentor, and develop the systems administration, lab operations, and technical services teams responsible for the facility supporting global Engineering and Research and Development. Own the reliability, efficiency, protection, safety, supportability, and continuous improvement of engineering labs, HPC systems, and infrastructure facilities. Establish service levels, operating standards, escalation paths, performance measures, monitoring, observability, automation, ticketing, and configuration-management practices. Translate engineering and customer requirements into infrastructure roadmaps, capacity p
Senior Technical Program Manager, Cloud Infrastructure — 2 Locations. Apply via Workday.
Senior Product Architect, K8s-Based AI Infrastructure — 2 Locations. Apply via Workday.
Equity Research Senior Associate, Communications & Infrastructure — New York New York United States. Apply via Workday.
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Software Engineers at Palantir drive large-scale transformation through data, AI and world-leading infrastructure that supports mission-critical workloads. In this role, you’ll have an opportunity to grow more quickly than you ever envisioned as you contribute high-quality code directly to: • Rubix and Apollo, platforms deployed at the most important institutions across the public and private sectors • Shaping Mission Manager, our new internal-infrastructure business line, used by advanced civil and defense agencies worldwide to power their infrastructure in highly sensitive environments • Building the core capabilities used by advanced civil and defense agencies worldwide to power their infrastructure • Providing the substrate on which Palantir deploys its other platforms, Foundry and Gotham, which power workflows for research scientists, aerospace engineers, intelligence analysts and economic forecasters You’ll join our Production Infrastructure organization, made up of small teams of engineers working on: • Environment Platform: a Kubernetes-based PaaS spanning hundreds of production clusters • Apollo: secure, fleet-wide deployment and change-management for complex microservice suites • Signals: our full suite of observability and alerting tools Core Responsibilities As a Software Engineer at Palantir, you’ll own every phase of the product lifecycle—from generating ideas and designing prototypes to executing features and shipping releases—while being paired with a dedicated mentor who champions your growth. You’ll work hand-in-hand with both technical and non-technical colleagues to uncover real customer problems and deliver solutions that ad
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime