Role Description As a Software Engineer on the Metadata team, you’ll build and operate the large-scale distributed databases that every Dropbox service depends on. Metadata systems are mission-critical, in the live path for all user operations and must meet stringent requirements for latency, durability, and transactional consistency. You’ll design and evolve the core infrastructure that manages Dropbox’s databases at scale, enabling fast, reliable access to data for millions of users and hundreds of internal services. This work spans distributed systems, replication, caching, and transactional database systems. You’ll collaborate closely with engineers across Infrastructure and Product teams to ensure the metadata layer meets business needs and continues to scale with Dropbox’s growth. This is an opportunity to leverage your expertise in distributed systems and grow into broader technical leadership. Our Engineering Career Framework is viewable by anyone outside the company and describes what’s expected for our engineers at each of our career levels. Check out our blog post on this topic and more here . Responsibilities Design and maintain distributed database systems providing low-latency, strongly consistent data access Implement and optimize replication, consensus, and caching mechanisms to meet availability and performance goals Operate production systems, including participating in the on-call rotation, ensuring high availability and data durability Collaborate with infrastructure and product teams to assess current and future use cases and requirements, supporting the development of a mid- to long-term roadmap that reflects these needs Contribute to system design reviews, postmortems, and reliability improvements Write high-quality, efficient code in Go and Rust for performance-critical systems On-call work may be necessary occasionally to help address bugs, outages, or other operational issues, with the goal of maintaining a stable and high-quality experienc
Jobiba hiring network
Infrastructure Team Manager Jobs
4,730 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Workforce Identity Cloud Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities, such as reducing costs and doing more for your customers. If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Staff Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimising costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh. Key Responsibilities: Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimised for production workloads, providing high resilience and operational efficiency. AWS Infrastructure Management: Build, manage, and optimise AWS cloud infrastructure, including EKS, ECS, S3, VPCS, RDS, IAM, and more. I
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
ABOUT THE TEAM We’re shaping the future of financial technology at Trendyol. As Trendyol’s technology teams, we’re not only building for today we’re designing the financial experiences of tomorrow. From payment infrastructure and digital wallets to smart credit systems and personalized financial services, we create solutions that empower millions of users across our ecosystem. With Trendyol Pay, we enable fast, secure, and seamless payment journeys. Through Trendyol Finance, we develop inclusive and accessible products that simplify financial decisions. We are united by a shared purpose:To create a positive impact in our ecosystem by enabling commerce through technology About the Role As an Application Security Engineer, you’ll work closely with Trendyol’s engineering teams to enhance application security across our platforms. You’ll perform web and mobile security testing, validate fixes through code reviews, and provide expert guidance on secure coding practices. In this role, you’ll support our bug bounty program, develop custom security tools, and help drive root-cause analysis and remediation. We’re looking for a collaborative, open-minded teammate who is eager to improve, communicate clearly, and promote security best practices throughout the organization.
About the Team OpenAI, in partnership with our capital and technology partners, is building a global network of advanced datacenters to support the most demanding AI workloads. The Industrial Compute team ensures that all datacenter systems are manufactured, delivered, and commissioned to the highest standards of quality, reliability, and performance. We work closely with manufacturing partners, engineering teams, and operations staff to ensure that every component is delivered ready for installation, startup, and long-term service. About the Role We are seeking an experienced Quality Engineer (QE) to drive Product and Site Quality initiatives across OpenAI’s infrastructure ecosystem. In this role, you will establish, implement, and manage a comprehensive, quality-focused program across our global supply chain network, ensuring excellence from design through deployment. You will be responsible for end-to-end quality of finished products, as well as maintaining and elevating manufacturing site quality standards. Working cross-functionally with Design (NPI) and Engineering teams, you will help achieve First Pass Yield (FPY), quality, and reliability targets. This includes leading site and fixture validation efforts, driving yield improvement initiatives (Yield Bridge, CPI), and implementing robust corrective and preventive actions (CAPA) to resolve issues at their root cause. In addition, you will play a key role in supplier quality management, assessing and qualifying new vendors, overseeing ongoing supplier performance, and ensuring readiness for future business awards. You will lead vendor audits, monitor key performance metrics, and coordinate corrective actions to ensure predictable delivery schedules, reduced operational risk, and high system reliability. By partnering closely with external suppliers and internal Engineering and Operations stakeholders, you will help ensure OpenAI’s datacenter infrastructure is delivered on time, meets the highest quality standa
About the Team pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model. About the Role We’re looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You’ll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production — directly improving how quickly and reliably research moves forward. In this role, you will: Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency. Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance. Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures. Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance. Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention. You might thrive in this role if you: Are excited about the potential of personal AGI and want to build the infrastructure that enables it. Have strong software engineering fundamentals and experience building or operating large-scal
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. At Stripe, our Mexico City office is a vibrant hub at the forefront of our mission to reshape the financial landscape for businesses worldwide. Our team prioritizes collaboration, innovation, and excellence. As a member of the Mexico City team, you'll be part of a mission-driven community dedicated to enhancing the global economy and increasing the GDP of the internet. We strive for excellence by creating with craft and beauty, while having fun and celebrating our successes together. We cultivate a culture of collaboration, inclusivity, and support where every team member's voice matters. Our commitment shines through as we handle over a million support cases each year, empowering our users not just to solve problems but to achieve their goals. About the team Stripe was built with simplicity in mind. We strive to deliver frictionless experiences for all of our users, whether they are an independent business, startup, SMB, or enterprise, and our mission is to provide all Stripe users with the best support experience possible. Today, Stripe handles over 1 million support cases per year and processes millions of internal transactions. We'll achieve excellence by thinking of support in a novel, solution-oriented way and viewing operations as an integral enabler of all our growth. Stripe has unique operational problems resulting from both our type of scale and the type of businesses we partner with as a result of "growing the GDP of the Internet." Stripes leverage underst
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. At Stripe, our Mexico City office is a vibrant hub at the forefront of our mission to reshape the financial landscape for businesses worldwide. Our team prioritizes collaboration, innovation, and excellence. As a member of the Mexico City team, you'll be part of a mission-driven community dedicated to enhancing the global economy and increasing the GDP of the internet. We strive for excellence by creating with craft and beauty, while having fun and celebrating our successes together. We cultivate a culture of collaboration, inclusivity, and support where every team member's voice matters. Our commitment shines through as we handle over a million support cases each year, empowering our users not just to solve problems but to achieve their goals. About the team Stripe was built with simplicity in mind. We strive to deliver frictionless experiences for all of our users, whether they are an independent business, startup, SMB, or enterprise, and our mission is to provide all Stripe users with the best support experience possible. Today, Stripe handles over 1 million support cases per year and processes millions of internal transactions. We'll achieve excellence by thinking of support in a novel, solution-oriented way and viewing operations as an integral enabler of all our growth. Stripe has unique operational problems resulting from both our type of scale and the type of businesses we partner with as a result of "growing the GDP of the Internet." Stripes leverage underst
About the Team OpenAI Finance is responsible for ensuring the organization is set up for success in pursuit of its mission. OpenAI’s Tax and Trade team sits at the center of OpenAI’s global growth—shaping how cutting-edge AI products, partnerships, and infrastructure scale across borders while navigating complex tax, trade, and regulatory regimes. We operate as strategic operators, not just compliance experts, embedding early in product, finance, policy, and infrastructure decisions to manage risk, unlock incentives, and enable OpenAI to grow responsibly and competitively worldwide. As OpenAI’s products, monetization models, and global infrastructure expand, our tax systems must scale with the same rigor and reliability as our product architecture. The team partners deeply with Financial Engineering, Product Engineering, and Finance Systems to build a modern tax technology stack that enables accurate, real-time tax determination across OpenAI’s global business. About the Role We're hiring a Director, Tax Infrastructure & Incentives to lead OpenAI's global tax infrastructure strategy supporting AI infrastructure expansion programs. You will own the tax strategy supporting data center development, site selection, infrastructure investments, and government incentive programs across. This role extends far beyond traditional property tax planning and reporting - you will help shape where and how OpenAI invests billions of dollars in AI infrastructure by partnering with executive leadership, infrastructure teams, governments, utilities, and external stakeholders. You will build scalable frameworks for negotiating and managing tax incentives, property tax, indirect tax, and infrastructure-related tax matters while developing repeatable processes that enable OpenAI to expand rapidly across multiple jurisdictions. You will also establish the governance, reporting, and operational infrastructure required to support long-term compliance, financial reporting, and executive
About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their in revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do You will act in a player-coach capacity and will be accountable for supporting associates in delivering against target set goals and mentoring in subject matter expertise. You will manage a small team of TechOps Payments Analysts. Responsibilities Champion metrics-driven analysis to continuously assess and optimize team performance, with a strong focus on automation Leverage automation to streamline data reporting and deliver timely, data-driven insights to management/key stakeholders Utilize data as a core tool for guiding people management strategies and optimizing team productivity Play a key role in weekly business reviews (WBR), using data to lead discussions support decision-making Manage capacity and scheduling, dividing and assigning work between team members. Lead independent discussions with TechOps, Engineering and XFN Partners to unblock complex reconciliation issues. Ensure your team has strong data analysis and Technical skills needed to be successful in their role. Setting clear goals and expectations for individual and team performance. Foster a culture of continuous improvement to refine team processes and procedures. Support recruitment and hiring initiatives. Coach and mentor individuals to meet career via structured career development conversations. Provide continuous performance feedback and facilitate periodic formal performance reviews. Drive and own initiatives that make the
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team Stripe MALPB operates in a highly complex, fast-moving regulatory environment. The Risk and Compliance team is responsible for ensuring that our expansion is grounded in a bulletproof governance framework that matches the speed of our technology. Having successfully established our foundational ERM policies, taxonomy, and risk registers, we are entering the integration phase. This team is building the future of "AI-Native" risk management—where governance frameworks are embedded into daily business operations through automated data pipelines, machine learning, and LLM orchestration. What you’ll do As the sole dedicated ERM Lead for Stripe MALPB, you will own the enterprise risk program end-to-end. This is a unique hybrid role designed for a technically fluent risk architect. Roughly 1/3rd of your time will be dedicated to ERM Integration and Execution: rolling out our mature risk framework to the First Line (1LoD), running the RCSA cycle, managing the control library, and synthesizing risk scorecards for executive and Board-level reporting. Roughly 2/3rds of your time will be dedicated to AI and Control Automation Design: acting as a hands-on developer to build and scale LLM tools, automated evidence-collection scripts, and real-time monitoring workflows that eliminate manual compliance overhead. Responsibilities ERM Framework Integration & Execution (1/3rd) Enterprise Rollout: Drive the cross-functional implementation and adoption o
About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. Our team builds the software, tooling, and operational systems that help manage this fleet at scale. We work across production engineering, distributed systems, capacity management, and operational automation to improve reliability, reduce manual work, and make better use of available compute. About the Role We are looking for a software engineer with experience building or operating large-scale production systems. You will develop the systems that help manage the GPU fleet powering ChatGPT, including tooling for fleet health, capacity planning, operational automation, and incident response. You will work closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and compute utilization. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. In This Role, You Will Build software and internal tools to manage large-scale GPU infrastructure supporting ChatGPT inference. Develop systems for capacity planning, fleet health monitoring, and resource utilization. Automate operational workflows, including incident detection, diagnosis, and response. Identify and address bottlenecks affecting fleet reliability, scalability, and performance. Partner with infrastructure, research, and product engineering teams to improve the compute platform. You Might Thrive in This Role If You Have experience operating large-scale production infrastructure, GPU clusters, or other compute-intensive distributed systems. Have a background in production engineering, site reliability engineering, infrastructure engineering, or platform engineering. Have built software that automates operational workflows and reduces manual work. Have worked with distributed infrastructure, cluster orchestration, or large-scale int
About the Team The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training. We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings. Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role As a Software Engineer on the Frontier Systems team focused on power management, you will work on critical infrastructure to support cutting-edge research. With large-scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. This role is critical to ensuring that our cutting-edge research supercomputing infrastructure runs smoothly, while maintaining reliability and grid-level power stability. Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change. This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions. We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role, you will: Develop and implement system-level and software-level solutions to optimize power usage in large-scale supercomputers, ensuring efficient and reliable operations. Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability. Work with researchers and engineers to design tools for real-time monitoring, detection, and remediation of power-related hardware and system faults. Collaborate cross-functionally to translate complex electrical system requirements into code, while driving continuous improvements in power man
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! The Data Platform team at Figma builds and operates the foundational systems that power analytics, AI/ML, and data-driven decision-making across the company. We serve a diverse set of stakeholders, including AI researchers, machine learning engineers, data scientists, product engineers, and business teams that rely on data for insights and strategy. Our team owns and scales critical platforms such as the Snowflake data warehouse, ML Datalake, orchestration and pipeline infrastructure, and large-scale data ingestion and processing systems, managing all data flowing into and out of these platforms. Despite being a small team, we take on high-scale, high-impact challenges. In the coming years, we're focused on building the data infrastructure layer for Figma's AI-powered products, driving cost and performance optimizations across our data stack, scaling our ingestion and reverse ETL capabilities for new product use cases, and strengthening data quality, reliability, and compliance at every layer. If you're passionate about building scalable, high-performance data platforms that empower teams across Figma, we'd love to hear from you! This is a full-time role that can be held from one of our US hubs or remotely in the United States. What you'll do at Figma: Design and build large-scale distributed data systems that power analytics, AI/ML, and business intelligence across Figma. Develop batch and streaming solutions to ensure data is reliable, efficient, and scalable across the company. Manage and evo
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Global Sales Enablement Team is responsible for partnering with cross-functional teams to equip our sales organizations with the skills, tools, and knowledge they need to perform at their best. We are currently leading one of the most significant transformations in Stripe's Sales history — reshaping how Sales Development operates through the adoption of AI-powered tooling and redesigning the SDR's role in driving pipeline quality. This transformation spans multiple geographies, sales motions, and hundreds of sellers. The stakes are high, the pace is fast, and the change management challenge is real. What you’ll do As a Sales Development Change Management Specialist, you will own the strategy and execution of our global change management and communications program for the Sales Development transformation. You sit at the intersection of strategy, content, and field engagement — responsible for ensuring that complex, high-priority initiatives are adopted at pace and land durably in the field. This is not a support role. You own the change narrative, the communication strategy, and the measurement of adoption. You'll partner with senior program leadership, subject matter experts, and regional enablement leads to drive consistent, high-impact transformation across a global sales organization. Responsibilities Own the end-to-end change management program — from preparing the organization for change through launch, adoption, reinforcement,
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime