About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing robust, high-performance systems that provide real-time insights, resource tracking, and seamless interaction with complex infrastructure. We aim to optimize resource allocation, minimize operational overhead, and create user-friendly tools that enhance researcher productivity and system transparency. About the Role You will design, develop, and operate web-based systems that provide a powerful and intuitive interface to OpenAI’s supercomputing clusters. You will collaborate closely with researcher, product and infrastructure teams to deliver scalable solutions that enable seamless monitoring, job scheduling, and resource management. This is an opportunity to work at the cutting edge of AI infrastructure, designing tools that scale to exascale workloads while maintaining usability and performance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and develop full-stack web applications to track, monitor, and manage large-scale AI workloads in real time. Collaborate with researchers and infrastructure teams to translate complex operational needs into intuitive UIs and scalable backends. Build data visualization tools (e.g., Gantt charts, dashboards) to provide insights into job scheduling and resource allocation. Optimize backend services to handle massive data throughput while ensuring low-latency performance and high availability. Implement frontend components that provide seamless interactions with scheduling, storage, and compute systems. Ensure system security, reliability, and scalability across globally distributed supercomputing infrastructure. You might thrive i
Jobiba hiring network
Infrastructure Team Manager Jobs
4,730 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil
About the Team OpenAI's Industrial Compute organization is responsible for planning, delivering, operating, and optimizing the compute infrastructure that powers frontier AI. As OpenAI scales toward becoming an intelligence utility, Industrial Compute coordinates a complex lifecycle spanning infrastructure strategy, capacity planning, provider partnerships, fleet operations, product demand, and financial planning. The organization manages one of the largest and fastest-growing compute footprints in the world, where decisions around capacity allocation, deployment readiness, utilization, reliability, and product demand directly impact product availability, customer experience, and business performance. The Capacity Systems team builds the software platforms, data systems, and automation frameworks that connect these functions into a shared operating model. We transform fragmented planning workflows into scalable systems that enable teams to understand what compute was contracted, delivered, healthy, allocated, and ultimately converted into business and research outcomes. About the Role We are seeking a Capacity Systems Software Engineer to build the platforms and services that power Industrial Compute planning, forecasting, optimization, and operational decision-making. In this role, you will design and develop software systems that connect infrastructure delivery, fleet health, capacity allocation, demand forecasting, deployment readiness, financial planning, and product consumption into a unified system of record. Your work will help OpenAI make better decisions about where compute should be deployed, how capacity should be allocated, and how infrastructure investments translate into business value. You will partner closely with Capacity Planning, Fleet Operations, Infrastructure Engineering, Product, Finance, Supply Chain, and Strategic Sourcing teams to replace spreadsheet-driven workflows with scalable software systems that enable visibility, automation, and dec
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Backend Engineer We’re redefining Privileged Access Management (PAM) from the ground up, purpose-built for Cloud, SaaS, Databases, Containers, and any virtualized environment. Our mission is to simplify and secure workforce access with seamless, secure-by-default workflows that adapt dynamically to modern infrastructure. We eliminate standing privileges, enforce least privilege, and embed Zero Trust principles into every access workflow by default. About the Role We are seeking a Staff Backend Engineer to serve as the core technical anchor and senior Individual Contributor (IC) for our newly established engineering pod in India. At the P4 level, your primary sphere of influence will be at the team level —taking ownership of complex, ambiguous problems and defining how to solve them cleanly, securely, and efficiently. In this role, you will lead by example through hands-on architecture, high-velocity coding, and end-to-end execution. You will drive the implementation of secure database and network device connectors (routers, switches, firewalls) on top of our core Zero Standing Privileges (ZSP) platform. You will work closely with our local Technical Team Lead to elevate the pod’s engineering craft, acting as a technical multiplier for mid-level developers while ensuring tight architectural alignment with our global team. What You’ll Be Doing Execution & Technical Impact End-to-End Ownership: Consistently design, code, debug, test, moni
Join the Team at Graphcore Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore brings together deep expertise to solve complex problems and deliver meaningful progress in AI compute. Ready to raise the standard for supplier quality in advanced AI manufacturing? Apply now to be part of the journey. Job Summary We are seeking a highly motivated individual with experience in managing supply chain quality in a high-tech manufacturing organisation. The ideal candidate will be hands-on, comfortable with ambiguity and able to manage multiple projects and stakeholders simultaneously. Working across our entire supply chain in a technically challenging, fast paced environment, you will ensure the readiness of our global supply chain to ramp successfully as we develop and manufacture the world’s most advanced AI systems and services. The Team The Graphcore Quality Team delivers outstanding customer experience and champions sustainable excellence throughout Graphcore. We are responsible for customer, supply chain and product quality, as well as organisational compliance, governance and assurance. You will be joining a diverse
Opportunity Overview: The Cohere IT team administers systems with a holistic approach and is responsible for the study, design, development, implementation, support, and management of telecommunications and computer-based information systems and software. Providing support to our colleagues who utilize Cohere Health systems in their day-to-day responsibilities and operations is a key function of the IT team.It is our aim to ensure a friendly working environment while scaling solutions to provide secure and efficient technology to all Cohere employees. IT is the backbone of our business, the underlying structure that reinforces Cohere’s growth, and propels this world-class organization to perform at the highest levels.The Cohere IT team is currently seeking an Information Systems Specialist to be part of the day-to-day frontline that provides, supports, and maintains the availability of services and infrastructure within the company’s corporate environment.Desktop Support Technicians are responsible for learning, maintaining, and documenting information pertaining to systems, software, applications, and processes used across the organization. All internal support requests are their responsibility to respond to or escalate to Systems Administrators, as necessary.Last but not least: The ideal candidate will thrive in a fast-paced startup with the ability to wear many hats while prioritizing customer focus. People who succeed here are empathetic teammates who are candid, kind, caring, and embody our core values and principles . We believe that diverse, inclusive teams make the most impactful work. Cohere is deeply invested in ensuring that we have a supportive, growth-oriented environment that works for everyone. What you'll do: Work within Cohere’s established standards, procedures and guidelines in delivering service to staff Fulfill service requests in a highly customer-service oriented manner Monitor, respond to, and escalate service issues in a timely manner
Who we are About the team Financial Connections is Stripe's open banking platform, enabling businesses to securely access consumer-permissioned financial data. Our platform connects to thousands of financial institutions, powering use cases from account verification to risk assessment to personal financial management. Across the Financial Connections Engineering org, we focus on delivering high-quality, enriched bank data at scale — building the ML systems that transform raw financial data into actionable signals for both internal Stripe teams and external merchants. Our ML work spans transaction categorization, risk scoring, data enrichment, and the development of intelligent systems that improve data quality across our network. We operate at the intersection of fintech infrastructure and applied machine learning, solving problems that directly impact Stripe's ability to serve millions of businesses and consumers. What you'll do We're looking for machine learning engineers who want to build intelligent systems that provide financial data at scale. You'll play a key role in designing, training, and deploying ML models that improve the quality, accuracy, and usefulness of financial data across Stripe's ecosystem. Responsibilities Design, build, train, evaluate, deploy, and own ML models in production that improve transaction categorization, risk scoring, and data enrichment across Financial Connections Design and build large-scale ML systems that operate on diverse financial data from thousands of institutions Experiment and iterate on ML models (using tools such as PyTorch, TensorFlow, XGBoost) to achieve key business goals around data quality and accuracy Develop pipelines and automated processes to train and evaluate models in offline and online environments Integrate ML models into production systems and ensure their scalability and reliability Collaborate with product, data science, and engineering partners across Stripe to identify opportunities where ML can im
NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and dedicated people on the planet working for us. If you're creative and autonomous, we want to hear from you! What you'll be doing: Developing software to enable efficient network design, deployment and day 2 management. Building product focused software solutions, used by internal and external customers. Helping us as we transform our workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies. Owning and driving integrations with various service APIs such as Cloud Service Providers, to automate creation of environments and auto populate data sources in turn. Building on open source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations. Streamlining deployment mechanisms and life cycle operations Developing modern service architectures around streaming data and event pipelines. Working with infrastructure domain experts on true, zero touch deployment solutions and utilizing best of breed high performance computing management solutions. Be a proactive problem solver, looking out for new opportunities to improve our services and customer experience. Communicate readily with your peers across the organization, b
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. We are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload. What you will be doing: Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure. Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration. Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks. Improve performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads. Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments. Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification. What we need to see: BS in Computer Science, Information Sys
The worldwide data management software market is massive. At MongoDB we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the forefront of innovation and creativity. MongoDB is seeking a Software Engineer 3 to join the Atlas Clusters Platform team. The team is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Clusters Platform team develops the foundational orchestration platform behind MongoDB Atlas. Our systems drive cluster planning and execution, evolve the Atlas control plane toward service-oriented architecture, and provide critical infrastructure that help Atlas run safely and efficiently across cloud environments. We are looking to speak to candidates who are based in New York City for our hybrid working model. What you’ll do Build and design new features for MongoDB Atlas Contribute to and lead complex technical projects Work closely with product and design teams, considering the user’s perspective while building technical solutions Work with customers and support engineers to fix issues Collaborate with team members to develop our codebase, best practices, and design principles Learn from and mentor other team members We’re looking for someone who Has at least 3 years of professional software development experience Is skilled at writing large-scale, distributed backend systems in a compiled language (Java, C#, Go, etc.) Is comfortable working across the stack of a modern web application (e.g. React, TypeScript, Kubernetes) Has experience with at least one major cloud provider technology (AWS, Azure, GCP) Has led the launch of a new feature and maintained it in production Is eager to solve tough problems Has excellent
About Supabase Supabase is the Postgres development platform, built by developers for developers. We provide a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. All services are deeply integrated and designed for growth. About the Role Supabase's Partnerships function is scaling fast - spanning Technology Partners, Solution Partners, Startups, and Cloud Partnerships. We're looking for a Partner Operations & Systems Lead to own the operational and technical backbone that lets this team move quickly and make decisions with good data. This is a senior, hands-on role. You'll design the systems, processes, and reporting that the whole partnerships org runs on - and, increasingly, you'll build the internal tools yourself. Supabase is investing heavily in AI-assisted development to move faster than a traditional ops build cycle allows, and this role is expected to be a leading example of that inside Partnerships. You'll report directly to the Head of Partnerships and sit alongside our regional and functional leads as a peer, with the mandate to build and enforce operational rigor across all teams. Beyond the infrastructure, we want someone who helps shape where Partnerships places its bets - not just the system that reports on them. What You'll Own Systems & Data Infrastructure Own the partner tech stack end-to-end - CRM partner objects, attribution tooling, and partner-facing portals. Design and maintain partner attribution models and the dashboards leadership uses to evaluate performance. Ensure data hygiene and consistency across partner records, deals, and touchpoints spanning all sub-functions. Process & Program Management Design and run partner onboarding, tiering, and certification programs. Own the RFC / DRI decision-making framework for the partnerships team, including how it's used and refined over time. Run deal-registration and "quarterback" account-ownership processes that keep internal teams
The worldwide data management software market is massive (IDC forecasts it to be $138 billion by 2026). At MongoDB, we are transforming industries and empowering developers to build amazing apps that people use every day. We are the leading modern data platform and the first database provider to IPO in over 20 years. Join our team and be at the center of innovation and creativity. MongoDB is seeking a Sr. Staff Software Engineer to join the Atlas Core Data Services organization. The organization is responsible for building MongoDB Atlas, our database as a service offering and fastest growing product, along with the API Platform and Developer Tools. Atlas allows users to deploy fault-tolerant, secure, globally distributed MongoDB clusters in just minutes. The Atlas Core Data Services organization builds the software that manages the Atlas cluster infrastructure hosted on the three major cloud providers (AWS, Azure, and GCP), as well as the software that manages the MongoDB database hosted on that infrastructure. We are constantly challenged to design features that ensure Atlas clusters are secure, available, durable, and performant while running large-scale, critical workloads. The Sr. Staff Engineer in this role will drive innovation across the organization and the company, setting technical standards and direction that enable future growth and velocity. We are looking for engineers with the experience and high standards needed to lead at that scale. Our organization champions a strong culture of inclusivity, diversity, and collaboration. If you want to be a deeply technical leader on a collaborative team that applies systems expertise to build the foundational infrastructure of a popular database, join us. Let's build a faster, more reliable, and highly scalable database platform together. We are looking to speak to candidates who are based in Dublin for our hybrid working model. Responsibilities Define standards and vision for the mission-critical Atlas SaaS data
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Staff Backend Engineer We’re redefining Privileged Access Management (PAM) from the ground up, purpose-built for Cloud, SaaS, Databases, Containers, and any virtualized environment. Our mission is to simplify and secure workforce access with seamless, secure-by-default workflows that adapt dynamically to modern infrastructure. We eliminate standing privileges, enforce least privilege, and embed Zero Trust principles into every access workflow by default. About the Role We are seeking a Staff Backend Engineer to serve as the core technical anchor and senior Individual Contributor (IC) for our newly established engineering pod in India. At the P4 level, your primary sphere of influence will be at the team level —taking ownership of complex, ambiguous problems and defining how to solve them cleanly, securely, and efficiently. In this role, you will lead by example through hands-on architecture, high-velocity coding, and end-to-end execution. You will drive the implementation of secure database and network device connectors (routers, switches, firewalls) on top of our core Zero Standing Privileges (ZSP) platform. You will work closely with our local Technical Team Lead to elevate the pod’s engineering craft, acting as a technical multiplier for mid-level developers while ensuring tight architectural alignment with our global team. What You’ll Be Doing Execution & Technical Impact End-to-End Ownership: Consistently design, code, debug, test, moni
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview The Ads DSE team builds the core capabilities that power Instacart's off-platform advertising, secure data collaboration (cleanrooms), automated taxonomy management, and their associated platform foundations. We are hiring a Senior Engineer II (L6) to co-own the technical roadmap, modernize our data solutions infrastructure, scale off-platform integrations and measurement capabilities, extend our cleanroom collaboration capabilities, and automate taxonomy change management workflows. This role is pivotal for cross-team collaboration, developer productivity, and AI-driven development initiatives. You are a senior technical leader who designs and ships scalable, data‑intensive systems; drives cross‑team execution; and raises engineering standards. You’ll partner closely with Product, Data Science, Data Platform, and Governance teams to build reusable abstrac
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime