Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Technical Program Manager for Compute Infrastructure, you will lead technical programs to develop and rollout next generation technology solutions to process a diverse spectrum of infrastructure workloads across Roblox services to enable users around the world to efficiently, reliably, and securely enjoy the Roblox Platform. You Have: Been a technical leader with domain expertise in systems and software used to process and manage Cloud and on-prem Infrastructure 7+ years of experience in software industry driving the build of large-scale infrastructure Experienced with establishing work relationships across multi-disciplinary teams and earning trust as a technical leader with all partners Experienced with identifying critical technical problems and opportunities Experienced with delivering end to end technical programs through roadmapping and reliable execution Able to turn specific solutions into systems that benefit the larger community in the long run Knowledgeable about user needs, scoping, planning, execution and delivery Flexible around process, using process as a tool when it makes sense for a team Experience with Kubernetes, Cloud platforms (e.g. AWS), and AI/ML infrastru
Jobiba hiring network
Lead Technical Program Manager Infrastructure Jobs
6,876 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current lead technical program manager infrastructure jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The mission of the Engineering TPM team is to drive Figma's most important cross-company engineering efforts, and we are looking for a Technical Program Manager (TPM) to partner with our Infrastructure team. The TPM provides oversight of the most important efforts that require coordinated technical execution across the Org to succeed. This is a role focused on enabling Figma's infrastructure teams to scale, improve performance, and deliver on critical projects. These large-scale efforts will involve collaboration across numerous backend, infrastructure, and security teams and cross-functional stakeholders, prioritization, decision-making, tracking execution, and driving operational excellence. We're looking for someone that can work in a TPM greenspace environment and is passionate about people, technology, and program management. Progress over process is our mantra. This is a full-time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Lead the execution, coordination, and risk management of Figma's infrastructure projects, ensuring seamless integration with minimal performance impact Drive key infrastructure initiatives, including reliability, storage, distributed systems, cloud-native perfo
Technical Program Manager – Applied Infrastructure About the Team The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety. The Applied Infrastructure TPM team partners across engineering to lead foundational programs that ensure OpenAI’s infrastructure can meet current and future demand. About the Role We’re looking for a seasoned Technical Program Manager to drive critical infrastructure programs across the Applied organization. This TPM will focus on cross-cutting initiatives such as general compute capacity planning, process transformation, cost and quota attribution and optimization, and coordination across infrastructure and product stakeholders. There will also be focus on evolving OpenAI’s infrastructure to support growth, scale and new products. This work is core to how OpenAI manages and grows its infrastructure footprint in a disciplined, scalable way. Location: San Francisco, CA (Hybrid – 3 days/week in-office) In this role, you will: Serve as the DRI for complex infrastructure programs spanning CPU planning, orchestration, and other resource management domains (e.g. networking, storage). Build and operationalize systems to capture demand signals, model future capacity needs, and align infrastructure planning across internal teams and partners external to the company. Partner closely with Infrastructure, Product and Finance teams to forecast infrastructure usage patterns and ensure supply/demand alignment. Lead cost attribution and quota enforcement programs to promote stability and ensure equitable access to resources across teams. Drive simplification and standardization of infrastructure tooling and processes across Applied and Infra organizations. Drive cross functional programs to evolve our infrastructure to support new growth and scale Work with external v
Location : Come and join us in Hamburg or Berlin! Following Lyft’s acquisition of Freenow, we are looking for a Senior TPM to lead the technical workstreams that get us to Day 1 readiness — the period between signing and closing — and to set up the integration that follows. This is a high-stakes, high-visibility role reporting into the EU Tech Leader, with direct exposure to Freenow and Lyft leadership. You will be the single point of accountability for technical Day 1 readiness. That means making sure every employee, every office, and every critical system is ready to operate underFreenow by Lyft from the moment the deal closes — and that the integration stream is up and running the day after. You will also be the TPM accountable for the integration of future acquisitions Freenow makes. This is not a role for someone who wants to administer a project plan. You are expected to be technically credible, opinionated, and comfortable holding senior people to account across IT, Security, Infrastructure, Engineering, Legal, People, and Finance. You make decisions when others stall, escalate cleanly when you can’t, and keep the whole machine moving. YOUR DAILY ADVENTURES WILL INCLUDE: Day 1 readiness. Own the end-to-end technical readiness plan between signing and closing. Define what “ready” means, track every workstream against it, and call the go/no-go on technical readiness for Day 1. IT and workplace cutover. Coordinate identity and system access (SSO, IdP, mail, productivity suite, engineering tooling), hardware provisioning and re-imaging, and office setup including network, VPN, Wi-Fi, meeting rooms, and physical access. Make sure nobody shows up on Day 1 unable to work. Security and compliance posture. Partner with Security and Legal to make sure the technical environment meets the new shareholder’s requirements on Day 1 — access controls, data handling, audit trails, vendor relationships. Integration stream kickoff. Stand up the technical integration workstreams
About the Team OpenAI's Enterprise team builds AI-powered enterprise products and shared platform capabilities that help organizations put advanced AI to work securely and at scale. Our work spans enterprise workflows, agent experiences, integrations, identity, administration, security, governance, and deployment. About the Role As a Technical Program Manager on Enterprise, you will lead the technical strategy and execution behind the products and shared capabilities that make ChatGPT, Codex, and future OpenAI products useful, secure, and scalable for organizations. You will translate customer needs, competitive dynamics, and product priorities into actionable plans, influence architectural direction, and deliver durable capabilities across application, platform, and infrastructure layers. The role requires deep technical fluency, strong product judgment, and the ability to move between hands-on execution and broader enterprise strategy. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Drive technical strategy and execution for enterprise product and AI workflow initiatives, from design through implementation, launch, customer rollout, and iteration. Partner with engineering teams to influence architectural direction, interface definitions, and implementation tradeoffs across full-stack products, APIs, integrations, and shared platform systems. Translate enterprise customer requirements into actionable product priorities across AI-powered workflows, agent experiences, integrations, permissions, data access, evaluations, identity, security, governance, and deployment readiness. Represent the needs of enterprise buyers, IT administrators, security teams, business leaders, developers, and end users in product and technical decisions. Identify adoption barriers, competitive gaps, and opportunities to make OpenAI products easier for organizations
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We lead complex technical programs that help Plaid scale its engineering platform. We partner across Engineering, Infrastructure, Data, Security, ML, Legal, and Product to deliver company-wide technical initiatives that improve reliability, scalability, and developer productivity. You'll lead strategic technical programs from planning through execution. You'll partner with engineering leaders to align stakeholders, manage dependencies, drive decisions, and ensure successful delivery of complex initiatives. You'll work across a variety of technical domains, adapting quickly to new challenges and helping teams execute effectively. As a Technical Program Manager, you will lead high-impact, cross-functional initiatives. As a generalist, you may work on a variety of programs. An example is one that strengthens Plaid's data and machine learning platforms. You will partner with engineering, product, data, legal, privacy, and business stakeholders to drive complex technical programs from planning through execution. Your work will help improve data governance, modernize machine learning infrastructure, and accelerate the adoption of trusted, high-quality datasets that power analytics, artificial intelligence
About the Team OpenAI's Enterprise team builds AI-powered enterprise products and shared platform capabilities that help organizations put advanced AI to work securely and at scale. Our work spans enterprise workflows, agent experiences, integrations, identity, administration, security, governance, and deployment. About the Role As a Technical Program Manager on Enterprise, you will lead the technical strategy and execution behind the products and shared capabilities that make ChatGPT, Codex, and future OpenAI products useful, secure, and scalable for organizations. You will translate customer needs, competitive dynamics, and product priorities into actionable plans, influence architectural direction, and deliver durable capabilities across application, platform, and infrastructure layers. The role requires deep technical fluency, strong product judgment, and the ability to move between hands-on execution and broader enterprise strategy. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Drive technical strategy and execution for enterprise product and AI workflow initiatives, from design through implementation, launch, customer rollout, and iteration. Partner with engineering teams to influence architectural direction, interface definitions, and implementation tradeoffs across full-stack products, APIs, integrations, and shared platform systems. Translate enterprise customer requirements into actionable product priorities across AI-powered workflows, agent experiences, integrations, permissions, data access, evaluations, identity, security, governance, and deployment readiness. Represent the needs of enterprise buyers, IT administrators, security teams, business leaders, developers, and end users in product and technical decisions. Identify adoption barriers, competitive gaps, and opportunities to make OpenAI products easier for organizations
About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa
Location Details: Canada - BC or ON (remote) At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join our team GoDaddy - Global Production Engineering looks after GoDaddy's global infrastructure, in the cloud and on-premises. We are hiring an experienced Technical Program Manager, focused on our AWS cloud infrastructure, to plan, lead and deliver complex cross-team initiatives. This is a heavily coordination-focused role: you will own the execution of a portfolio of AWS cloud platform and cost-savings programs, working hands-on with software engineers, engineering managers and SREs to achieve outcomes aligned with the strategy. You will drive dependencies end-to-end, facilitate trade-off decisions, and give collaborators and leadership clear, reliable access to status and risk. You will be an integral part of the Technical Program Management team, partnering closely with engineering leads to ensure GoDaddy delivers on its planned objectives and global strategy. What you'll get to do... Own end-to-end delivery of a portfolio of concurrent cloud platform programs, coordinating across engineering and partner teams to manage scope, schedule and dependencies against the critical path. Drive cost-savings program coordination, including tracking, reporting and surfacing risks to goals and achievements proactively. Run intake and prioritization processes and keep priority pages and status sources current and trustworthy. Serve as the central coordination point across teams, facilitating trade-off and negotiation discussions, driving alignment, and resolving roadblocks with minimal issues. Build reports, scorecards and dashboards to c
About the Team The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role You will be part of an engineer-first TPM team as a Technical Program Manager for Compute Infrastructure who owns the end-to-end delivery of large-scale GPU clusters, partnering with engineers to bring clusters online across external providers and partners. You’ll run a broad, parallel portfolio spanning hardware, networking, power, and cooling—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference. Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths. Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams. Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale. Partner with engineering to improve cluster turn-up reliability, repeatability, and automation
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role We're looking for a Technical Program Manager who can operate at the intersection of engineering, product, and business — someone deeply technical, trusted instinctively by engineers, and sharp enough to drive clarity and momentum across complex, cross-functional programs. This is a high-agency role with real executive visibility and direct impact on how Ramp's engineering organization scales. We're looking for someone who is energized by complexity, deeply curious about what AI can unlock for engineering teams, and eager to apply it hands-on in their work. You should be someone who experiments with AI tools regularly, thinks about how they change the way software gets built, and brings that perspective into how you run programs. What You’ll Do Lead large-scale technical programs across engineering and adjacent teams—from CI/CD and infrastructure scaling to incident response, and driving other strategic projects across the engineering organization Own Ramp's engineering incident response program, improving processes, running retrospec
About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? As Cohere continues to grow, the size and complexity of our programs has continued increasing over time. To help us manage this, we are looking to bring in exceptional Technical Program Managers (TPM) to manage this! At Cohere, TPMs are pivotal to our success, being seen as operational experts specializing in cross-functional efforts. Great TPMs at Cohere are well-rounded, executing with precision at the tactical level, while partnering with senior leaders and shaping the big picture at a strategic level, and the opportunities are endless in growing the scale and impact of their work. Cohere’s TPMs are often described by their stakeholders as organized, execution-oriented, pragmatic, and adaptable to the needs of the company and the programs they lead. As a reward for the depth and breadth of expertize that you bring to the table, you will get to work alongside some of the most talented engineers in the world on building cutting edge AI technology. No day will be the same as the one before, and the dynamic high growth environment will be absolutely perfect for anyone excited about solving novel problems, bringing
About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req
The Anyscale Technical Program Management (TPM) team is expected to play a critical role executing high impact programs while continuously improving processes to sustainably grow and increase the effectiveness of the Tech organization spanning Design, Engineering and Product teams. As a Technical Program Manager focused on Anyscale’s core product solution, you’ll help lead complex application development in service of enhancing our product platform. In this role, you'll support and help scale the technical solutions that make Anyscale’s products and services possible. As part of the overall development life cycle you’ll plan requirements, identify risks, manage schedules, and communicate clearly with project stakeholders on complex projects with significant bottom line impact. Your Program management contributions will span prioritization, planning of projects and features, stakeholder management, tracking of external commitments while contributing to the organization's technical culture by highlighting and espousing best practices. You’ll learn and grow alongside talented teammates who share your commitment to excellence and appetite for innovative problem-solving. This is a rare opportunity to join in the leadership of a team that will be responsible for building a successful commercial ML/AI oriented solution from the ground up! As part of this role, you will: Help us build, track and ship our Commercial / OSS product Will work closely with the software development and product teams to deliver high quality, scalable products used by customers around the world Collaborate with the product teams and align all the stakeholders to assemble project teams, assign responsibilities, identify appropriate resources needed, and develop schedules to ensure timely completion of projects by meeting project milestones Assess risks, anticipate bottlenecks, provide escalation management, make tradeoffs, balance the business needs versus technical constraints and encourage risk ta
Get new lead technical program manager infrastructure jobs by email
Daily job updates · Unsubscribe anytime