About the Team OpenAI’s Legal team plays a crucial role in advancing our mission by tackling innovative and fundamental legal issues in AI. The team includes professionals from diverse legal fields—technology, AI, infrastructure, privacy, IP, corporate, employment, tax, regulatory, and litigation—who collaborate closely with colleagues across the company. If you are passionate about being a technology lawyer working on cutting-edge challenges, you’ll thrive here. About the Role We’re seeking a senior lawyer to participate in commercial legal strategy and execution across OpenAI’s fast-growing infrastructure portfolio. This is a cross-functional role that will partner closely with procurement, supply chain, partnerships, finance, and product teams to structure, negotiate, and manage the transactions that will support OpenAI’s long-term infrastructure ambitions. We’re looking for an experienced infrastructure transactions lawyer who thrives in ambiguity and wants to help define the commercial playbook for infrastructure efforts in the AI era. This role reports to the Associate General Counsel for infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3-days in the office per week and offer relocation assistance to new employees. In this role, you will: Own commercial legal strategy and risk management for OpenAI infrastructure transactions. Draft, negotiate, and advise on complex agreements with infrastructure suppliers, manufacturers, distributors, and technology partners. Support strategic partnerships involving AI infrastructure and hardware supply chains. Develop frameworks for procurement, licensing, and collaboration across the infrastructure ecosystem. Partner with finance and operations teams to align contract terms with business and compliance requirements. Collaborate with policy and regulatory colleagues on issues impacting global supply chains, export controls, and manufacturing. Build scalable, efficient contracting process
Jobiba hiring network
Infrastructure Team Manager Jobs
4,730 active opportunities · Updated for October 2026
Fresh results
15 shown
Explore current infrastructure team manager jobs. Use filters to narrow by work mode, employment type, experience and date posted.
The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency. Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT. We prioritize safety, reliability, and responsible AI deployment over unchecked growth. About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure. You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and build systems to manage both cloud and bare-metal fleets at scale. Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms. Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows. Automate infrastructure processes, reducing repetitive toil and improving system reliability. Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack. Continuously improve tools, automation, processes, and documentation to enhance operational efficiency. You might thrive in this role if you: Have strong software engineering skills with experience in large-scale infrastructure environments. Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers). Have deep expertise in server-level systems (e.g., systems, containerization, Chef,
About the Team The Statsig team within OpenAI owns the experimentation, rollout, dynamic configuration, and analytics infrastructure that sits on the launch path for OpenAI products. Our systems help teams ship safely, evaluate product and model changes in production, and make high-confidence decisions from real-world usage. Statsig began as an independent company built around experimentation, feature management, and product analytics at scale. After Statsig joined OpenAI, the team began the next chapter: bringing that platform expertise and infrastructure into OpenAI as the experimentation and rollout foundation for every product we ship. This is infrastructure with a very direct product consequence. Teams working on ChatGPT, Codex, model measurement, consumer experiences including ads, business subscriptions, developer products, and shared platform systems depend on Statsig to evaluate configurations, move traffic safely, ingest experiment data, serve analytics, and roll changes forward or back when production reality demands it. We are at a critical point in the platform journey. Adoption is accelerating quickly across OpenAI, and the systems that were already important are becoming load-bearing for how the company launches. The infrastructure needs to stay fast under sharply increasing evaluation volume, reliable when more services depend on it, observable enough to debug quickly, and efficient enough to support OpenAI-wide scale. Recent SDK and server-side infrastructure work has already produced measurable wins in latency, reliability, memory usage, and compute efficiency across important services. The next phase is to make those gains systematic: a platform that can absorb rapidly growing product velocity while preserving low latency, data quality, operational safety, and developer trust. Based out of OpenAI's Bellevue office, we are a close-knit team that values in-person collaboration, technical depth, operational ownership, and building infrastructure that
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Stripe Delivery Center (SDC) strategy will provide operational leverage and expand Stripe’s portfolio of operational capabilities to support the scaled needs for external users and internal Stripe teams. Stripe’s People & Places Operations team will sit within the SDC and is a highly collaborative, cross-functional team that drives key support for the People and Places function. We work with many different teams at Stripe. We are nimble and flexible individuals that can wear many different hats. We don’t mind working through ambiguity and love adding organization to chaos. We believe that success is not defined by any one individual, but rather by the collective work of the entire team What you’ll do Responsibilities Provide strategic and operational leadership for the newly-formed People Support team in CDMX, translating the global Core Ops strategy into an effective local operating model. Lead and develop a small team People Support specialists through clear expectations, regular coaching, actionable feedback, and quality calibration. Establish local team practices, coverage, escalation paths, and ways of working that integrate effectively with the global People Support organization. Manage daily operations, including queue health, workload allocation, service levels, and escalations, while contributing directly during the team’s ramp or periods of high volume. Use service metrics, case trends, and feedback to set priorities, i
JOB TITLE Access and Identity Management - Business Analyst A CAREER WITH POINT72'S TECHNOLOGY TEAM As Point72 reimagines the future of investing, our Technology group is constantly improving our company’s IT infrastructure, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts experimenting, discovering new ways to harness the power of open source and AI solutions, and embracing enterprise agile methodology. We encourage professional development to ensure you bring innovative ideas to our products while satisfying your own intellectual curiosity. WHAT YOU'LL DO Translate business and security requirements into scalable, production‑ready AIM and IGA solutions that support firmwide operations Partner closely with business, Global Systems Administration, and technology stakeholders to deliver secure, efficient, and user‑friendly access solutions Analyze data to identify access trends, control gaps, and process inefficiencies, with a focus on automating manual processes and improving platform scalability Own requirements gathering, functional documentation, and user stories from ideation through UAT, deployment, and post‑release support Leverage tools such as SailPoint, JIRA, SQL, Splunk, and Power BI to support access governance, reporting, and operational insights Design and maintain dashboards and KPIs that measure access lifecycle performance, risk indicators, and user experience Support data‑driven decision making by validating data quality, ensuring consistency, and aligning metrics to business objectives Build deep domain expertise in identity governance, security controls, and investment‑driven technology platforms Contribute to a culture that prioritizes integrity, transparency, and the highest ethical standards WHAT’S REQUIRED 4+ years of experience as a Business Analyst supporting Access & Identity Management (AIM) and Identity Governance & Administration (IGA) initiatives
About the Role: The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming. We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. This is a hybrid role for our Toronto office. What You'll Do: Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. Support best practices in experimentation, evaluation, and ML system monitoring. Ensure cost efficiency, scalability, and performance in ML infrastructure investments. Your Background: 10+ years of industry experience spanning machine learning engineering and distributed systems. 3+ years of leadership and management experience, with a proven ability to build and lead strong t
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The credit risk operations team plays a critical role in ensuring a healthy financial ecosystem for businesses around the world. This team will directly impact the company's bottom line and growth capabilities by supporting new and emerging businesses. The Credit Operations team is responsible for conducting credit risk assessments, underwriting small and mid-size users, and proposing mitigation strategies that enable Stripe to take smart risks for our users. What you’ll do In this role, you will manage and develop a group of Credit Risk Operations Associates that are focused on commercial credit underwriting and risk mitigation. The Credit Risk Operations Team Lead will cultivate the engagement of their team members while guiding them to be the best they can be, through feedback, coaching, mentoring, and advocacy within the organization. This means helping to set team goals, and using metrics to efficiently measure and guide team performance in pursuit of those goals. To be a fit, you will have a strong operations mindset, possess a deep understanding of commercial credit risk, be able to move quickly, and be passionate about delivering an incredible user experience. Responsibilities Manage, coach, and develop a new team of in-office Credit Risk Operations Associates Execute on commercial credit underwriting frameworks that allow your team to evaluate business models and financial statements of successful venture-backed startups to esta
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We’re entering a world where AI agents don’t just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents’ ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova’s throughput, correctness, and operational rigor grows dramatically. We’re looking for a Staff Software Engineer who wants to go deep on both the engine internals and the infrastructure underneath it. You’ll work across the full stack of a modern OLAP system — query planning and execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — while driving meaningful improvements to performance, cost-efficiency, and reliability at scale. You’ll influence technical direction through your work, your design reviews, and your mentorship of other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of customers. What You’ll Do Build and evolve core query engine infrastructure Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Design for high-throughput automated quer
About the Role & Team Every AI insight, every experiment, every cohort at Amplitude starts with a query. Our in-house OLAP engine, Nova , processes trillions of events in real time — turning raw behavioral data into fast, trustworthy answers that power decisions for thousands of product teams worldwide. We're entering a world where AI agents don't just assist product teams — they ship features, run experiments, and make prioritization calls autonomously. What makes that possible is agents' ability to verify their work against real product data continuously. That makes Nova the critical infrastructure in the loop, and as non-stop agents become the main source of queries, the demand on Nova's throughput, correctness, and operational rigor grows dramatically. We're looking for a Senior Software Engineer who wants to go deep on the engine internals and the infrastructure underneath. You'll own significant components of a modern OLAP system — across query execution, columnar storage and encoding, distributed compute, caching, and cloud infrastructure — and drive meaningful improvements to performance, cost-efficiency, and reliability. You'll grow your technical influence through the quality of your code, your design contributions, and your collaboration with other engineers on a team of ~10. This role is ideal for someone who finds real satisfaction in making a complex distributed system faster, cheaper, and more reliable — and who wants to do that work on a system that directly powers the product experience for thousands of enterprise customers. What You'll Do Build and improve core query engine components Contribute across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. Implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. Help ensure Nova's components support
About the job Who are we? About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About The Team The Stripe Delivery Center (SDC) strategy will provide operational leverage and expand Stripe’s portfolio of operational capabilities to support the scaled needs for external users and internal Stripe teams. Stripe’s People & Places Operations team will sit within the SDC and is a highly collaborative, cross-functional team that drives key support for the People and Places function. We work with many different teams at Stripe. We are nimble and flexible individuals that can wear many different hats. We don’t mind working through ambiguity and love adding organization to chaos. We believe that success is not defined by any one individual, but rather by the collective work of the entire team. Key Responsibilities Provide strategic HR leadership to HR Operations. Effectively manage, develop and engage the global team Ensure an exceptional employee experience by simplifying key processes Lead and implement HR initiatives and projects which are aligned within HR and Centers of Expertise (COE) Coordinate and lead key projects for improvement across HR Identify best practices that can be applied to improve work tasks and processes Deliver service improvement activity across HR through employing process improvement methodologies and the application of innovative thinking Promote and lead change Supports the administration and maintenance of HR systems while driving process and data integrity across the HR landscape P
Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team The Payouts and FX team helps global businesses manage their money wherever it flows. We own Stripe's FX and multi-currency management products—including FX Acquiring, Multicurrency Settlement (MCS), and Currency Conversion—which together serve tens of thousands of global merchants and platforms. Our team works at the intersection of payments, treasury, and financial infrastructure, helping businesses reduce FX costs, eliminate third-party dependencies, and consolidate their multi-currency operations onto Stripe. What you'll do In this role, you will own the full commercial cycle—prospecting, structuring, closing, and activating deals—across Stripe's FX Acquiring, Multicurrency Settlement (MCS), and Currency Conversion (ICC) products. You will engage CFO- and Treasurer-level stakeholders at global merchants and platforms, positioning Stripe's multi-currency suite as the complete alternative to standalone FX management providers. You will also serve as a product specialist for the broader Account Executive team, building the playbooks, deal structures, and GTM motion that scale this business. Responsibilities Own quota and commercial responsibility across FX Acquiring, Multicurrency Settlement, and Currency Conversion, prospecting and closing deals with global merchants and platforms Structure and negotiate holistic multi-currency deals that bundle MCS and Currency Conversion, pricing across both products to deliver merchant value while captu
Join the Atlas Search team to design and develop the next generation of Semantic and Vector Search infrastructure. Atlas Search is a growing cloud service that allows users to execute complex search and vector search queries using the MongoDB Query Language. Our users can focus on relevance and data retrieval instead of the machinery needed to search data at scale. Our team is building a cloud-based distributed system responsible for the core components of search including data ingestion, performance, query language, query execution, for both relevance-based search and vector search. Our product is being adopted quickly and there are many interesting projects. This is a technical role where you will be responsible for the infrastructure and features enabling our at-scale cloud service powering vector and semantic search. We are looking to speak to candidates who are based in the San Francisco Bay Area for our hybrid working model. What You’ll Do Lead complex projects across the MongoDB ecosystem, for instance, development of a new Search deployment framework within the MongoDB managed cloud Set project level strategy, architect features, and lead projects to successful execution Identify, design, and implement features enhancing our reliability, performance, security and efficiency Perform code reviews with peers and make recommendations on how to improve our software development processes Influence and grow team members through active mentoring and leading by example What We Look For 5+ years experience in data management/search systems, ideally with a strong distributed systems and infrastructure background Experienced in the development and maintenance of concurrent, stateful services Eager to shape the technological direction of a complex system and have the ability to lead initiatives through collaboration with others Experienced in writing features, debugging and optimizing multithreaded applications written in Java Familiarity with LLM
Location: Remote (U.S.) Clearance: U.S. Citizen or Permanent Resident Required with the ability to obtain a Public Trust Salary Range: 110K – 125K LTS is seeking a Senior Security RMF Engineer to join a cybersecurity transformation surge team supporting the VA.gov Platform. This role will serve as the bridge between VA security/RMF requirements and the engineers responsible for implementing those requirements across VA.gov. The Senior Security / RMF Engineer must understand how security controls are implemented in modern cloud infrastructure and software delivery environments and be able to translate control deficiencies, authorization requirements, and security risks into actionable engineering work.This is not intended to be a documentation-only compliance role. This individual will work closely with DevSecOps engineers and the existing VA.gov Platform ATO/security team to assess the current security posture, address gaps in VA.gov's Critical Controls, support ATO/cATO readiness, improve authorization artifacts, and automate evidence and control assessment wherever possible. The PWS specifically describes the desired model as one in which ATO/RMF documentation confirms security rather than defines it, with success measured through risk reduction and security outcomes rather than paperwork completeness. What You’ll Do: Assess VA.gov Platform compliance with the 18 Critical Controls identified by VA and help establish a baseline of current implementation and remaining gaps. Perform security reviews, gap analyses, and risk assessments across VA.gov Platform infrastructure, pipelines, applications, and component systems. Support ongoing ATO and cATO readiness for the VA.gov Platform authorization boundary. Develop, update, and maintain RMF and authorization artifacts, including System Security Plans (SSPs), control narratives, POA&Ms, Business Impact Analyses (BIAs), Privacy Threshold Analyses (PTAs), and supporting evidence. Evaluate identified
About Remote Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. We make it possible for businesses of all sizes to recruit, pay, and manage international teams. With our core values at heart and future focused work culture, our team works tirelessly on ambitious problems, asynchronously, around the world. You can find Remoters working from 6 different continents (Antarctica left to go!) and all of our positions are fully remote. With Innovation as one of the core values, we have built Automation and AI capabilities into the requirements for every role. We encourage every member of the Remote team to bring their talents, experiences and culture to the table to help us build the best-in-class HR platform. If you are energetic, curious, motivated and ambitious, be part of our world. Apply now and define the future of work! What this job can offer you Remote's SRE team exists so that our engineers can move quickly and our customers get a product that stays up. The team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, our observability stack and the reliability practices that sits on top of all of it. We are looking for a Team Leader to run that team. This is a 60% IC, 40% leadership role. You will own the career development of your reports, steer the teams focus using judgment against the company goals, and you will be the spokesperson for the team across engineering. You will also stay close enough to the technical work to set direction with credibility and to know when something is going wrong before it is escalated to you. Reliability practice at Remote is maturing rather than mature. Our SLO framework is live on its first few teams and needs to reach the rest, there is real work to do on how we balance operational load against project delivery. If you want a team where the foundations are in place and the interesting problems are still open, this is that team. What you bring People leadership Yo
Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment. Our customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast-growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security. Backed by world-class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people-centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI-ready operations. Forward is currently seeking experienced Java developers to work as part of our Network team. Responsibilities Help bring the best ideas from the software development world into the networking industry. Contribute to our code base, systems and software architecture as a member of our engineering team. Help create and optimize network device models for different device vendors and protocols. Help create infrastructure needed to configure, collect and test network devices. Work with peers who are experts in Networking, Distributed Systems, Big Data and Search. Requirements 5+ years of work experience in software development 3+ years of work experience with Java BS in Computer Science or related degree Solid software engineering experience with large code bases Basic understanding of networking and TCP/IP. Strong verbal and written communication skills. Nice to haves Working knowledge of how switches, routers, firewalls or load balancers work. Experience working with networking protocols such as BGP/OSPF/IS-IS, IPv4/IPv6, MPLS, VLAN, VXLAN, etc. This position is a re
Get new infrastructure team manager jobs by email
Daily job updates · Unsubscribe anytime