We are seeking an experienced IT/Lab Manager to lead the planning, deployment, and operations of our physical lab environment and IT systems. This role will focus on building and maintaining scalable, reliable, and secure environments to support engineering teams involved in research, quality assurance, validation, and related activities. It will also support internal collaborators. You will have an outstanding opportunity to drive innovation in a multidimensional, technology-focused company that is crafting the future of data-center and lab technologies. If you bring perfection and creative thinking while solving issues as they arise, and enjoy working with distributed teams – your place is with us! What You’ll Be Doing: Own day-to-day operations, planning, and roadmap for the engineering lab and IT infrastructure (servers, storage, networking, and related services). Lead and mentor an IT/Lab team, driving guidelines, standards, and a culture of ownership, partnership, and continuous improvement. Collaborate closely with R&D, QE, Verification, and other engineering teams to design, provision, and maintain environments that meet their performance, reliability, and security needs. Lead all aspects of running data center and lab operations, including rack layout, cabling, power and cooling, hardware lifecycle, and resource availability. Lead procurement and vendor management for hardware, software, and services, including evaluation, negotiation, and ongoing relationship management. Implement and maintain automation for system provisioning, configuration, and operations using tools such as shell/Perl/Ansible. Design and maintain monitoring, logging, and alerting for servers, network, and storage systems to ensure high availability and rapid incident response. Investigate and resolve sophisticated infrastructure issues across OS, networking, storage, virtualization, and appli
Jobiba hiring network
Infrastructure Team Manager Jobs
4,815 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses. We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey. We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do. And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora. Become an Engineering Manager / Senior Engineering Manager at Bloomreach and lead our Infrastructure team as a core foundation for product engineering. In this role, you will own the platform capabilities that power personalization across our clients’ websites and mobile apps, enabling fast experimentation and multi-variant testing at scale. You will lead the team responsible for our core infrastructure stack - Google Cloud Platform (GCP), databases, observability platform, and Kubernetes, and partner closely with Product, Security, and application engineering teams to ensure our platform is reliable, scalable, secure, cost-efficient, and developer-friendly. Your leadership and strategic direction will impact hundreds of millions of end customers across diverse e-commerce verticals. This is a full-tim
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About the Team The ChatGPT Search Product Infrastructure team builds the foundational systems that power search experiences across ChatGPT. We develop the product infrastructure that connects models with search systems and other sources of real-time information, enabling ChatGPT to deliver timely, relevant, and trustworthy answers to users around the world. Our work sits at the intersection of product engineering, AI, and large-scale infrastructure. We build shared platforms and abstractions that enable product teams to independently develop, evaluate, and launch new search-powered experiences. These platforms provide the guardrails, testing capabilities, observability, and rollout controls needed to prevent reliability, scalability, quality, and latency regressions while supporting rapid product iteration. The team partners closely with: Post-Training on model launches, experimentation, and prompt optimization Search product verticals on new user experiences Inference on GPU efficiencies Indexing and Retrieval on the systems that identify and deliver relevant information Capacity/Fleet team to ensure optimal regionalized provisioning of GPUs and CPUs About the Role We are looking for an Engineering Manager to lead the team responsible for ChatGPT’s Search Product Infrastructure. You will set the technical and organizational direction for the systems that bring search capabilities into ChatGPT. You will guide architectural decisions across search orchestration, model and prompt integration, serving infrastructure, experimentation, observability, evaluation, and product integrations. You will balance immediate launch and product needs with the long-term reliability, scalability, latency, and maintainability of the platform. A central responsibility of this role is creating leverage for Search product verticals. You will lead the development of extensible platforms that allow those teams to independently build, test, and launch features without requiring ongoing invol
The Development Infrastructure team builds the infrastructure for the development platform our Asana engineers use every day. We build and operate the software that accelerates how Asana R&D delivers new products to our users. We combine industry best-practices and innovation to enable Asana to remain product-forward delivering a high-quality beloved experience for our users. You will focus on developing the team around you, helping them contribute more and grow as both engineers and individuals. You will channel your passion and enthusiasm as much into recruiting and building a team, as the technical challenges your team will be resolving. This role is based in our Reykjavik office with an office-centric hybrid schedule. The standard in-office days are Monday, Tuesday, and Thursday. Most Asanas have the option to work from home on Wednesdays. Working from home on Fridays depends on the type of work you do and the teams with which you partner. If you're interviewing for this role, your recruiter will share more about the in-office requirements. What you’ll achieve Manage and lead a team of 3-7 engineers, including a technical lead, with focus on impactful deliveries Define how the team evolves and interacts with the rest of Asana Recruit for and grow the team Lead the team to identify future opportunities to improve development velocity and the developer experience Influence the future of software development at Asana and guide the technical direction across the organization Collaborate with a world-class team of engineers to enable new development capabilities through infrastructure About you 3+ year of engineering management experience The rare mix of intelligence, empathy, integrity and technical skills that allow you to rapidly earn the trust of a technically astute team Focus on maximizing impact, for yourself and your team Ability to provide valuable input to any technical or product discussion Oriented around the multi-year consequences
About the Role & Team We’re looking for an Engineering Manager to lead the Data Infrastructure team within Statsig Experiment at Amplitude. You will lead a multidisciplinary team of software engineers, data engineers, and data scientists responsible for the systems that power experimentation at scale. The team owns three critical areas: Data ingestion: Collecting and importing experiment exposures, custom events, OpenTelemetry data, and real user monitoring data across SDKs, streaming systems, cloud storage, and customer data warehouses. Data computation: Building distributed computation systems that transform raw data into accurate, timely experiment results. Stats engine: Developing and productionizing the statistical methods that help customers make trustworthy decisions from their experiments. This is not a traditional data engineering management role. We are looking for a leader with a solid data science and statistical foundation who can connect advances in experimentation methodology with scalable production systems. You will help set our technical and scientific direction, translating new statistical methods and machine learning research into capabilities that customers can use reliably at scale. You’ll partner closely with data scientists, engineers, product managers, and customers to advance the state of experimentation. The ideal candidate is equally comfortable discussing causal inference and statistical power with data scientists, distributed computation architectures with engineers, and experimentation strategy with customers. What You’ll Do Lead and grow the team responsible for Statsig’s data ingestion, experiment computation, and stats engine. Define the technical and scientific strategy for advancing experimentation across both Statsig Cloud and warehouse-native deployments. Partner with data scientists and engineers to turn new statistical and causal inference methods into scalable, reliable product capabilities. Evolve our data and computatio
Squarespace provides innovative solutions to empower our customers to focus on building their brand and growing their businesses on our platform. The Databases team manages all of the backend infrastructure that Squarespace runs on – MongoDB, CockroachDB, and Kafka clusters, to name a few examples. We are an accomplished, diverse group of people who develop the services that guarantee reliable and scalable infrastructure for both our cross-functional partners in product engineering, as well as our end users on the Squarespace platform. We believe that infrastructure excellence doesn't stop at just building for today; it needs to have a solid foundation of scalability, reliability, and a robust developer experience for the future. This is a hybrid role working from our Dublin office 3 days per week. You will report to the Databases Senior Engineering Manager. You’ll Get To… Nurture high-performing software engineers by guiding navigation when there is ambiguity. Distill the scope of the team and help hire a balanced group of engineers that will excel as a unit. Grow the career development of direct reports through regular 1:1s with direct, actionable feedback. Celebrate wins that motivate the team’s positive culture and robust dynamic. Evaluate consistently to improve team efficiency and effectiveness when required. Evolve a deep understanding of local systems to identify appropriate architectural decisions. Thread with Product, Design & Engineering to champion, define and execute an optimal roadmap. Bond across Engineering, Product, Design, Marketing, Data Science and Business Operations. Who We’re Looking For 3+ years of recent experience managing a Product Engineering team of four or more engineers. 7+ years of industry experience deploying apps across large codebases with many contributors. Ability to fluently translate, document and present technical concepts to non-technical stakeholders. Strong technical foundations to navigate the inherent tra
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The mission of the Engineering TPM team is to drive Figma's most important cross-company engineering efforts, and we are looking for a Technical Program Manager (TPM) to partner with our Infrastructure team. The TPM provides oversight of the most important efforts that require coordinated technical execution across the Org to succeed. This is a role focused on enabling Figma's infrastructure teams to scale, improve performance, and deliver on critical projects. These large-scale efforts will involve collaboration across numerous backend, infrastructure, and security teams and cross-functional stakeholders, prioritization, decision-making, tracking execution, and driving operational excellence. We're looking for someone that can work in a TPM greenspace environment and is passionate about people, technology, and program management. Progress over process is our mantra. This is a full-time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Lead the execution, coordination, and risk management of Figma's infrastructure projects, ensuring seamless integration with minimal performance impact Drive key infrastructure initiatives, including reliability, storage, distributed systems, cloud-native perfo
NVIDIA Networking is a leading provider of innovative end-to-end InfiniBand and Ethernet connectivity solutions for servers and storage. Our portfolio includes adapter cards, switches, cables, and software designed to optimize Data Center performance with industry-leading bandwidth and scalability. We serve diverse sectors such as high-performance computing, enterprise, cloud computing, and Web 2.0. Our mission is to stay ahead of the market by delivering groundbreaking products and services. Our Ethernet solutions are tailored for industries like Media & Entertainment and any domain that benefits from advanced DataStream and TCP/IP acceleration. What You’ll Be Doing: Lead a team of 8+ mechanical design engineers. Define priorities, create project plans, and allocate resources for mechanical programs in coordination with Product Managers. Drive all electro-mechanical, automated JIG and thermal design aspects, of production test setups, ensure readiness of test and assembly infrastructure for high-volume manufacturing. Develop multiple early design concepts in fast-paced product development cycles. Lead task forces to investigate and resolve production issues, reliability concerns, and conduct failure analysis. Perform risk assessments and implement mitigation strategies during product design. What We Need to See: B.Sc. in Mechanical Engineering or higher. 10+ overall years of relevant experience including 4+ years of experience managing teams of engineers in R&D environment. Proven expertise in developing, testing, and manufacturing of complex automated connection systems with precise moving parts, pneumatic and electro-mechanical systems design. Strong hands-on experience with 3D CAD tools (Creo preferred) static and dynamic mechanical simulation Solid background in designing components and sub systems an
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Developer Infrastructure team within Coinbase's Platform product group builds the tools and platforms that every Coinbase engineer relies on to build, test, ship, and operate software. As the Group Product Manager for this team, you'll own the product vision and multi-year strategy for the entire code lifecycle, from code change to production deployment. You'll define what world-class developer infrastructure looks like at Coinbase and drive the execution that makes every engineer faster, safer, and more productive. What you’ll do: Own the end-to-end product roadmap and success metrics for Developer Infrastructure, including CI/CD pipelines, release automation, testing infrastructure, deployment systems, and production readiness standards. Drive a continuous simplification agenda by leading regular systems migrations that reduce complexity, improve reliability, and make the platform easier for engineers to adopt and use effectively. Build quality scorecards and leaderboards for production systems across Coinbase, ensuring all systems meet defined standards with executive-level visibility and accountability. Partner with Engineering, SRE, Security, and infrastructure leadership to ensure developer lifecycle systems are reliable, performant, secure, and scalable. Deepen understanding of internal customer needs by analyzing usage patterns, gathering feedback f
Ready to do the most impactful work of your career? At Coinbase , we are uncompromising on our mission to increase economic freedom. The bar is high, the environment is intense, and we like it that way. This isn't a place for complacency, it’s a place to be pushed past your perceived limits. If you're ready to build the future of finance alongside people who refuse to settle for "good enough," you belong here. Coinbase is a remote-first, but not remote-only company. Expect to get together quarterly for intense in-person working sessions called “surges.” learn more about working at Coinbase . The Core Infrastructure team within Coinbase's Platform product group builds the foundational systems that keep Coinbase online, secure, and scalable, owning the compute and networking platforms that power every product and service across the company. As the Group Product Manager for Core Infrastructure & Reliability, you'll own the product vision and multi-year strategy for Coinbase's cloud infrastructure, driving the design, operation, and scaling of the systems that underpin hundreds of billions of dollars in annual transaction volume. You'll partner deeply with Engineering, SRE, Security, and Finance to ensure Coinbase's infrastructure is reliable, cost-efficient, and resilient across multiple cloud environments and regions. What you’ll do: Own the product strategy and roadmap for Core Infrastructure, spanning compute, networking, multi-region and multi-cloud architecture, and platform reliability. Strengthen infrastructure reliability and resilience programs, defining platform-level SLOs, capacity planning, failover capabilities, and incident reduction targets to meet the uptime demands of a global financial platform. Lead evaluation and adoption of cloud infrastructure technologies (Kubernetes, service mesh, distributed storage, observability, infrastructure-as-code), making build-vs-buy decisions that balance cost, speed, and long-term scalability. A
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The infrastructure teams provide efficient and optimized infrastructure for Stripe to build secure, reliable, and differentiated products, while enabling Stripe developers to achieve their highest potential. Stripe makes it easy for any developer to access and manage the capabilities of the financial system while maintaining the least regulatory friction. We work to enable developers to have the most productive results of their entire career from the very first days they join Stripe through years of developing new systems and products. What you'll do As a Technical Program Manager in the Infrastructure team, you'll play a key role within engineering and drive programs that span across Stripe in the core infrastructure of Stripe's payment systems. You're responsible for the successful definition, cross-functional strategy, planning, and execution of large-scale technical programs that help to solve complex problems and enable products and infrastructure at scale. You'll deliver outstanding results by building and implementing solutions that scale while protecting our users and serving the business. Responsibilities Work with teams across the organization to understand pain points in their infrastructure usage to find common ideas and work to create solutions that span multiple domains. Define and produce high-quality written proposals, communications, and documentation. Execute on technical programs that require deep systems and engineering-l
Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role: The Web Infrastructure team builds the foundations behind Notion’s web clients, including client architecture, performance, reliability, and shared design systems. As Engineering Manager, you’ll lead the team through the evolution from Notion Clients. You’ll set strategy, develop senior engineers and managers, and partner across Notion to make the product faster, more reliable, and easier to build. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days. What You'll Achieve: You'll build and manage a diverse and inclusive team of engineers and managers working on core parts of Notion's architecture. You'll create a healthy environment in your team that embodies Notion's values. You'll recruit, coach, and develop engineers. You'll ensure engineers are regularly receiving feedback and making progress on personal and professional goals. You'll facilitate planning—the prioritization, sequencing, and staffing of work—for your team. You'll be responsible
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Engineering Manager, Home Infrastructure The Home Infrastructure team builds the mission-critical backend and data systems that power Roblox’s Homepage and Experience Details Page, two of the highest-traffic surfaces on Roblox. These surfaces reach the vast majority of Roblox’s daily active users and are core drivers of discovery, engagement, retention, and platform growth. We are a full-stack product infrastructure team responsible for content distribution across Roblox. Our systems support multiple modes of user interaction, including exploratory browsing, directed discovery, and personalized content recommendations across the many types of content that make up the Roblox ecosystem. This team sits at the intersection of large-scale distributed systems, machine learning-powered personalization, data infrastructure, and product experimentation. We partner closely with Machine Learning, Data Science, Product, Design, Frontend, Ads, Marketplace, Virtual Economy, and other teams across Roblox to build the platforms that help users find the most relevant and engaging content. As Engineering Manager for Home Infrastructure, you will lead a team of Backend and Data Engineers responsible for the e
Fin is the AI Customer Agent company on a mission to help businesses provide perfect customer experiences. Our AI Agent Fin is the highest-performing AI Customer Agent on the market today, enabling businesses to deliver impeccable, always-on customer support across the customer journey – from service, to sales, to ecommerce. Powered by our own AI models, Fin resolves complex customer issues end-to-end across every channel, with minimal set-up and integration. Fin can also be combined with our natively integrated Intercom help desk for one single system that is designed to meet the needs of modern day support teams. Founded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Driven by our core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. Fin is the AI Customer Service company on a mission to help businesses provide incredible customer experiences. Our AI agent Fin, the most advanced customer service AI agent on the market, lets businesses deliver always-on, impeccable customer service and ultimately transform their customer experiences for the better. Fin can also be combined with our Helpdesk to become a complete solution called the Fin Customer Service Suite, which provides AI enhanced support for the more complex or high touch queries that require a human agent. Founded in 2011 and trusted by over 30,000 global businesses, Fin is setting the new standard for customer service. Driven by our core values, we push boundaries, build with speed and intensity, and consistently deliver incredible value to our customers. What’s the opportunity? We’re hiring an Engineering Manager for our AI Models Infrastructure Team in the AI Group . The AI Models Infrastructure team builds and operates the foundational infrastructure th
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime