Jobiba hiring network

Infrastructure Team Manager Jobs

4,815 active opportunities · Updated for October 2026

Fresh results

15 shown

Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.

O
Okta
📍 Washington• Full-time• From $204K/yr
1mo ago

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. **This position requires 2 days a week in our San Francisco Office. The IDaaS Site Reliability Engineering Group Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling. As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling. What you’ll be doing Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform. Drive

awskubernetesci/cd
View job →
S
Stripe
📍 San Francisco• Full-time• Remote
1mo ago

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team You will be joining Stripe’s ML Foundations and Gen AI team to incubate new ML applications and improve our ML capabilities across Stripe. Our team is responsible for unlocking novel ML and LLM techniques and applications across Stripe’s product suite to drive business outcomes, as well as providing infrastructure, tooling and support for ML teams. What you’ll do As a senior product leader, you will lead a cross-functional team to define, incubate and scale new ML/AI applications across Stripe’s product suite, and drive our strategy and roadmap for ML/AI infrastructure powering all of Stripe’s teams. You will work closely with product leaders across business units to define and deliver on an AI-centric product strategy, launching new applications that drive incremental business outcomes. At the same time, you will be advancing our core AI technology stack to empower teams across Stripe to infuse their scenarios with Agents and agentic capabilities, with API support for agent quality and continuous improvement. Responsibilities Develop and execute on the Stripe-wide strategy for new ML/AI applications across our product suite Evaluate and align on areas of investment for ML/AI applications in collaboration with product leaders across the company Work with cross-functional teams to execute on the roadmap and launch successful new ML/AI applications Communicate clearly and crisply with leadership stakeholders and drive alignment acr

REMOTEairust
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Through strategic partnerships and self-built campuses, we are scaling one of the world's fastest-growing AI infrastructure platforms. The Supply Chain organization ensures critical infrastructure components—from compute systems and networking equipment to integrated rack solutions—are sourced, manufactured, qualified, and delivered with the speed and reliability required to support frontier AI development. We partner closely with Hardware Engineering, Manufacturing Quality Engineering, Infrastructure Delivery, Hardware Operations, Finance, and suppliers worldwide to build a resilient, scalable supply chain capable of supporting rapid infrastructure expansion. As Industrial Compute continues to grow, Supply Chain serves as the operational bridge between engineering innovation and large-scale infrastructure deployment. About the Role We are seeking a Supply Chain Manager to lead strategic execution across sourcing, supplier operations, manufacturing quality, and infrastructure delivery for OpenAI's AI infrastructure portfolio. This role will oversee a multidisciplinary team responsible for strategic sourcing, manufacturing quality engineering, and technical program management while partnering closely with engineering, finance, hardware operations, and deployment teams. You will drive supplier strategy, manufacturing readiness, production planning, quality performance, and operational execution across the full hardware lifecycle. Success requires balancing long-term supplier strategy with day-to-day execution. You'll establish scalable operating mechanisms, strengthen supplier partnerships, manage complex cross-functional programs, and ensure OpenAI can rapidly deploy AI infrastructure without compromising quality, cost, or reliability. This is a people leadership role responsible for developing a high-performing organization while driving operati

awsrestai
View job →
O
24 days ago

About the Team OpenAI’s Industrial Compute organization is building the infrastructure required to develop and operate increasingly capable AI systems at global scale. The Manufacturing Operations team works across Hardware Engineering, Manufacturing Engineering, Manufacturing Quality, Rack Integration, System Enablement, Supply Chain, Logistics, Deployment, and Hardware Operations to convert complex hardware designs into reliable production systems. We partner closely with ODMs, JDMs, contract manufacturers, and component suppliers to ensure that servers, racks, networking equipment, and supporting infrastructure are manufactured, validated, and delivered at the quality and scale required by OpenAI. About the Role We are seeking a Technical Program Manager to lead manufacturing operations programs across OpenAI’s hardware supply base in Singapore and the broader APAC region. You will own cross-functional execution from new product introduction through production ramp and sustaining operations. You will coordinate manufacturing partners and internal engineering teams around factory readiness, capacity, material availability, build plans, validation, quality gates, issue resolution, and delivery commitments. This role requires strong technical fluency, disciplined program management, and the ability to operate directly with manufacturing partners in fast-moving, high-stakes environments. You should be comfortable working at both the factory floor and executive-review levels, translating complex manufacturing conditions into clear risks, decisions, and recovery plans. This role is based in Singapore and requires regular travel to manufacturing partners across the APAC region. Key Responsibilities Lead manufacturing operations programs for AI servers, racks, networking systems, and related infrastructure across regional manufacturing partners. Own integrated program plans spanning NPI, factory readiness, material availability, tooling, test development, qualification,

REMOTEpythonsqlaws
View job →
PE
Private Employer
📍 Bangalore, Karnataka• Full-time
1mo ago

About Meesho Meesho is India's fastest-growing internet commerce company, on a mission to democratize e-commerce for everyone. We serve millions of customers and over 1.75 million sellers through technology-driven innovation, building the scalable systems that power Meesho's most critical surfaces — Search, Recommendations, Personalized Ranking, Logistics, Fraud Detection, and Image Match. The AI Platform sits at the heart of this. It serves a peak of 1M+ real-time deep-learning model inferences per second on ordinary days, scaling 3x+ on sale days — with the reliability that scale demands. The team works at the frontier of applied AI and infrastructure — multi-region inference, novel embedding-search algorithms, and optimized open-weight LLM models — squeezing out every bit of computation and passing the cost savings straight back to customers. About the Role We are looking for an experienced Engineering Manager – AI Engineering to lead the development of scalable AI platforms and infrastructure while managing high-performing engineering teams. You will drive the design, delivery, and optimization of production-grade AI systems powering AI use cases across Meesho.

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Digital Platform Team The Digital Platform team is a global organization distributed across the US and India, responsible for the integration and infrastructure layer that connects Adobe Experience Cloud to the rest of Okta's marketing technology, and analytics stack. The team owns Adobe Experience Platform (AEP) architecture, data collection infrastructure, AEM platform services, and the integration portfolio — including the MCP server architecture, LLM provider integrations, and AI agent infrastructure that power AI-native marketing experiences across the company. The team partners closely with Marketing, IT, DevSecOps, and Adobe to extend platform capabilities, govern data, and deliver scalable AI-driven experiences to the GTM organization. The work is strategic and product-driven: anticipating capability gaps, evaluating new tools and integrations, and building the connective tissue that keeps the broader ecosystem unified and well-governed

reactawsgit
View job →
O
1mo ago

About the Team OpenAI builds powerful AI systems like ChatGPT, the OpenAI API, and enterprise products that serve millions of users across the globe. As we scale, securing our infrastructure, protecting sensitive data, and meeting global compliance standards are essential to our success and societal impact. Security at OpenAI is a cross-cutting function that spans infrastructure, applied engineering, legal, policy, and product. Technical Program Managers (TPMs) play a critical leadership role in aligning teams and delivering execution at scale and this role will be foundational in shaping how we secure OpenAI’s systems, users, and commitments. About the Role We’re seeking a Senior Technical Program Manager to drive cross-functional security, privacy, IT and compliance initiatives at the intersection of infrastructure, product, and policy. You will execute complex programs that reduce internal data access, prevent misuse, and ship security capabilities. You will focus your efforts on the most critical initiatives within Security, crossing the spectrum of insider threat, information security, physical security, and information technology challenges. This role is deeply technical and execution-focused. It requires a structured operator who thrives in ambiguity, partners effectively across boundaries, and applies principled judgment to scale trust, governance, and security across OpenAI’s systems and products. In this role, you will: Drive execution of critical security and compliance programs such as vulnerability management, merger and acquisition security and integration, infrastructure hardening, and datacenter security management. You will need to deeply collaborate on technical architecture and resolve technical problems in partnership with engineering. Partner with IT, Infrastructure, Application, Legal, Privacy, and Security teams to build scalable programs, and deliver critical security outcomes across multiple disciplines, including insider threat, information

awsrestai
View job →
O
OpenAI
📍 United States• Full-time
1mo ago

About the Team The Stargate organization is responsible for building and scaling the physical infrastructure systems that power OpenAI’s next generation of AI training and inference platforms. This includes the manufacturing, deployment, and operational execution required to bring large-scale compute infrastructure online globally. The team operates at the intersection of data center infrastructure, hardware manufacturing, supply chain, deployment operations, and systems planning. We partner closely across Infrastructure Strategy, Manufacturing Operations, Capacity Planning, Supply Chain, Deployment, and Engineering to execute one of the largest infrastructure scale-outs in the industry. About the Role We are seeking a Technical Program Manager, Rack Delivery to drive operational execution across rack manufacturing, site readiness, and deployment coordination for Stargate infrastructure programs. This role will serve as a key connective layer between manufacturing partners, deployment teams, and infrastructure readiness programs to ensure rack production and delivery timelines remain aligned with site availability and deployment sequencing. You will help manage operational execution across contract manufacturers (CMs), support build planning and RCCA processes, and coordinate deployment readiness across multiple concurrent infrastructure programs. You will also partner closely with Demand Planning teams to translate strategic planning inputs into actionable SKU-level manufacturing and delivery schedules. This role is ideal for someone who thrives operating across ambiguity, manufacturing operations, infrastructure deployment, and large-scale cross-functional execution. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation support. Key Responsibilities Drive cross-functional coordination between rack manufacturing, deployment operations, and site readiness programs. Manage operational execution acros

awsrestai
View job →
O
OpenAI
📍 San Francisco• Full-time
1mo ago

About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st

awsazurerest
View job →
O
1mo ago

About the Team OpenAI Consumer Devices is building the next generation of products that bring powerful AI into people’s everyday lives. Guided by OpenAI’s mission to ensure AGI benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating useful, intuitive, and responsible technology. You’ll have the opportunity to work alongside exceptional people on ambitious, zero-to-one challenges at the intersection of hardware, software, and AI. This is a chance to help define an entirely new category of products—and shape how people experience AI in the future. The Systems Integration team is critical in this mission, turning complex hardware-software development into reliable product signals. We build the shared infrastructure, tooling, and lab environments that let teams test quickly, understand failures, and ship with confidence. About the Role As a Systems Integration Manager , you will lead the team responsible for device validation infrastructure, test automation, developer tooling, and lab operations. This is a player-coach leadership role: you’ll set technical and operational direction, build and develop a team of engineers and lab operations professionals, and stay close to the architecture and hardest systems problems. You will partner closely with device software, OS, firmware, hardware, reliability, QA, and release infrastructure teams to define validation strategy, improve release readiness, and ensure our test environments and quality signals scale with the product. Because this is a new category of devices, you’ll have the rare opportunity to build the validation foundation early—shaping the systems, standards, and operating model that will support products from prototype through launch. We’re looking for a leader who combines strong technical judgment with people leadership, operational rigor, and experience building reliable systems for complex hardware-software products. This role is b

artificial intelligenceai
View job →
M
Mongodb
📍 San Francisco• Full-time
1mo ago

Join our MongoDB Search Systems Engineering team at the forefront of building next-generation database infrastructure at scale. You'll work alongside some of our most talented and deeply technical engineering teams who are architecting distributed systems that power highly performant search and AI applications for enterprise customers worldwide. This team operates at the intersection of infrastructure innovation and customer impact—building the foundational platform capabilities that will define how organizations leverage search and AI at scale for years to come. In this role, you'll serve as both strategic partner and technical translator, helping brilliant engineers navigate the path to operational excellence while maintaining the forward vision to anticipate market needs 18-24 months ahead of current development cycles. What Makes This Role Unique: This is about architecting the future of search infrastructure. You'll be working with engineers who live and breathe distributed systems, helping them channel that brilliance toward platforms that don't just work today, but anticipate the architectural needs of tomorrow's AI-native applications. If you light up at the challenge of seeing around corners in infrastructure and can hold your own in technical debates about consensus protocols while keeping the team focused on customer value—this is your role. This role can be based out of our San Francisco office or remotely in USA region. Role Responsibilities Contribute to the Future of Search Infrastructure: Drive the long-term technical vision for MongoDB Search, anticipating emerging patterns in vector, hybrid, and AI-native applications to build conviction around platform investments 12–18 months ahead of market demand Champion the Customer & Innovation: Act as the voice of customers building mission-critical AI applications, using deep technical discovery to uncover latent needs around scalability and consistency, ensuring our architecture leads the industry in

reactmongodbaws
View job →

About the Team Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5.5! The Data Movement area owns the infrastructure that keeps training jobs supplied with the right data at the right time, and keeps model state moving safely and efficiently across large clusters. Our work spans machine learning systems, distributed storage, high-throughput data loading, reliability engineering, and developer experience. Success means researchers can move quickly while training runs remain fast, reproducible, debuggable, and resilient at scale. About the Role We are looking for a deeply hands-on Technical Lead Manager to own datasets throughout our training infrastructure. This person will set the direction for how training jobs read data: the APIs, storage contracts, versioning model, benchmarks, debugging tools, and reliability guarantees that make data access consistent across current and future training frameworks. You will begin as the primary technical owner for dataset reads, working directly in the code while aligning researchers, training framework owners, storage teams, and infrastructure partners around a durable platform. The problem is deceptively hard at frontier scale: make enormous, heterogeneous datasets easy to consume, correct across distributed workers, observable when something goes wrong, and flexible enough to support pretraining, reinforcement learning, and multimodal training. In this role, you will Design and build a unified dataset read platform for multiple current and future training frameworks. Define dataset APIs, storage-format expectations, registration/versioning, and migration paths that make data access reproducible and maintainable. Build reliability into the read path, including stateful iteration, caching, fast restart, recovery, and clear operational contracts. Build terminal and web-based visualizers that let teams inspect text, multimodal, and reinforcement learning data late

pythonawsrest
View job →

About Team Our Robotics team is focused on unlocking general-purpose robotics and advancing toward AGI-level intelligence in dynamic, real-world environments. Working across the full model and systems stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the physical constraints of real-world systems to improve people’s lives. About the Role We’re looking for a Technical Program Manager to own and scale the systems that power robotic data acquisition across our development and evaluation environments. This role sits at the intersection of Robotics Engineering, Operations, and Infrastructure, ensuring that DAQ stations and associated workflows reliably produce high-quality data for model training and evaluation. You will drive end-to-end execution of complex, cross-functional programs that integrate robotic platforms, sensing systems, operator tooling, and data pipelines into cohesive, production-ready systems. Success in this role requires strong systems thinking, operational rigor, and the ability to translate ambiguous research needs into scalable infrastructure. In this role you will: Own DAQ Program Delivery Lead the roadmap, execution, and scaling of robotic DAQ systems, ensuring alignment with research, engineering, and operational priorities. Drive Cross-Functional Integration Coordinate across robotics hardware, software, infrastructure, and operations teams to deliver tightly integrated, deployment-ready data collection systems. Operationalize Data Collection Systems Translate experimental and research workflows into repeatable, scalable DAQ processes with clear SLAs, metrics, and reliability targets. System Readiness & Deployment Ensure DAQ stations (robots, sensors, compute, operator interfaces) are fully integrated, validated, and ready for production use across multiple sites. Program Execution & Risk Management Build and manage detai

awsrestai
View job →
N
12 days ago

About the Team Netlify is a product-led growth company, and Revenue Operations owns the systems, data, and infrastructure that turn that growth into revenue. With a large volume of people discovering and signing up for Netlify every day, Marketing Operations is critical to understanding those users, connecting activity across our systems, and creating scalable ways to engage them throughout their journey. As Senior Marketing Operations Manager, you’ll own the operational foundation behind that work. You’ll connect HubSpot, Salesforce, product data, and analytics, build the systems and lifecycle infrastructure that turn signals into action, and give Marketing and our GTM partners trusted visibility into performance. Reporting to the Senior Director of Revenue Operations, you’ll translate the Marketing Operations vision into a roadmap and independently drive it forward while partnering closely with Growth, RevOps, Data, Product, and other GTM teams. What You’ll Do Execute on the Marketing Operations roadmap, translating broader Marketing and GTM priorities into scalable systems, processes, and programs. Own our marketing automation platform end to end, including workflows, lifecycle stages, segmentation, forms, lists, email operations, data hygiene, deliverability, governance, and documentation. Build and maintain the connection between marketing automation, Salesforce, product data, and analytics so we can understand and engage users throughout the customer journey. Establish reliable attribution and measurement across marketing and product-led journeys, including UTM governance and connections between campaigns, signups, activation, conversion, and revenue. Build reporting and self-service dashboards that give Marketing and GTM partners a trusted view of performance and support weekly, monthly, and executive-level decision-making. Partner with Growth to operationalize lifecycle programs across email, in-product experiences, and sales-assist triggers, turning c

REMOTEaisalesforceCRM
View job →
O
1mo ago

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT engineering org builds and operates the systems that bring product improvements to users across backend services, web, mobile, and desktop platforms. The Developer Velocity team partners with product engineering, platform, infrastructure, reliability, engineering acceleration, and observability teams to make everyday development faster and releases safer, more predictable, and easier to operate. About the Role We are seeking a Technical Program Manager to improve developer velocity and deployment excellence across ChatGPT. You will lead durable improvements to local development, CI, testing, build systems, release trains, progressive rollout, and post-deployment validation. You will identify the highest-leverage sources of engineering friction, align teams around shared standards and metrics, and drive adoption of tooling and workflows that improve both speed and reliability. This role combines systems thinking, technical program leadership, and hands-on operating rigor across a broad engineering surface. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the cross-functional roadmap for improving local development, CI, testing, build workflows, and release infrastructure. Create a durable intake and prioritization mechanism for developer friction, using evidence to focus teams on the highest-impact improvements. Lead programs that improve deployment speed and safety, including pre-merge confidence, progressive rollout, release guardrails, rollback readiness, and post-deploy valida

awsci/cdrest
View job →
🔔

Get new infrastructure team manager jobs by email

Daily job updates · Unsubscribe anytime