Jobs in United States

Infrastructure Sourcing Operations Lead in United States

1,486 active opportunities · Updated October 2026

Explore current infrastructure sourcing operations lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.6%

$155K – $400K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Events Analytics Platform (EAP) team is responsible for the infrastructure that powers all of Sentry's time-series data and searching capabilities across billions of events with sub-second latency. We started this initiative by building Snuba, the primary storage and query service for Sentry's event data powered by ClickHouse, and we are now focused on unlocking deeper visibility and reporting across the terabytes of event data our users generate. As a Senior Software Engineer, you will lead efforts to push the boundaries of data visibility at Sentry. You will do this by expanding the capabilities of our search infrastructure, building new capabilities on top of our state-of-the-art storage layer and increasing the performance and integrity of Sentry’s core data services. You will also help shape Infrastructure's technical direction at Sentry and collaborate with Product and other Engineering teams to turn that vision into a reality. If you want to solve the hard problems that come with scaling event data into the petabyte range, this could be the job for you. In this role you will: Expand EAP's ability to deliver data at world-class speed and reliability. Architect and automate services and systems to scale reliably under growing demand. Make architectural trade-offs that balance product requirements with engineering constraints. Maintain and grow the team's code quality initiatives by regularly reviewing code and contributing to design decisions. Lead design and discussions around deliverables the team is working towards. Improve the maintainability and developer experience of the codebases EAP owns. Exa

PythonSQLPostgreSQLRedis
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -73.5%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The FinOps function is responsible for financial accountability, visibility, and optimization across all engineering-related spend at Plaid. This includes cloud infrastructure, AI/ML and data workloads, third-party SaaS tools, and other technical investments that support Plaid’s products and internal platforms. The team operates at the intersection of Engineering, Product, and Finance, ensuring that spending decisions are transparent, intentional, and aligned with product strategy and business priorities. Rather than functioning as a cost-control or approval layer, FinOps enables teams to understand, own, and optimize their spend while maintaining engineering velocity. Responsibilities Monitors and analyzes engineering spend across cloud, AI/ML, data platforms, and SaaS, identifying trends, anomalies, and optimization opportunities. Builds and maintains forecasts for engineering spend, partnering with Finance and engineering leaders to understand drivers, assumptions, and risks. Partners with engineering, product, and TPMs to incorporate cost considerations into roadmaps, architectural decisions, and execution plans. Leads cost optimization initiatives, such as rightsizing, commitment strategies, an

SQLAWSAzureGCP
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -86.8%

From $187K/yr

Quick readStrong listing-quality and freshness signals

As a Cloud Security Engineer you will partner with different stakeholders across the organization to secure our cloud infrastructure. As part of the Platform Security organization we secure the building blocks of Datadog’s applications and infrastructure. We do this by building solutions to solve systemic risks and combine an approach of making the secure path easier and the insecure path harder to secure and accelerate the business. We regularly partner with the most bleeding edge internal products and are working to solve and build solutions to enable our safe usage of AI. We also develop AI based solutions to enable security at scale. We are looking for a Service Mesh and Kubernetes focused security specialist to help round out an incredibly strong infrastructure security focused group. You will rotate through a variety of internal projects and gain deep exposure to Datadog’s infrastructure. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Solve our most challenging cloud infrastructure security problems starting with our core building blocks and golden paths. Enable our engineers to build and ship secure solutions quickly. Build and extend Datadog’s Platform Security solutions. Leverage and influence the direction of Datadog’s products to secure our infrastructure, and provide internal feedback that enables our teams to improve the products for ourselves and our customers. Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or related scientific field or equivalent professional experience. Passionate about advocating for and implementing solutions to complex problems, at-scale, in a large multi-cloud environment. You don’t want to just provide security recommendations, you want to help imple

PythonAWSAzureGCP
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -97.1%

From $127K/yr

Quick readStrong listing-quality and freshness signals

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

MongoDBAWSAzureGCP
S
📍 United States· Full-time
✓ Quality checkedCompany trend -94.6%

Partner Solutions Architect, Technology Partners Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. About the team As a member of the Partner Solutions Architecture (PSA) team, you'll be responsible for providing technical guidance to Stripe's technology partners. Our goal is to drive growth through co-solutioning, education, and enablement with our partner organization. The right candidate will be a technical liaison to the Stripe Partner Development Manager. This will involve creating technical business plans for technology partners, including identifying, incubating, and bringing to market service and solution offerings built on or with Stripe. A Partner Solutions Architect is someone who is intellectually curious, open, and entrepreneurial. What you'll do As a member of the team, PSAs build consultative relationships with executives and technical decision-makers in the customer's organization to provide business and technical thought leadership, and become a trusted advisor in solving business-critical challenges. We architect business models with them so they can drive new monetization opportunities and growth. Responsibilities Work with key technology partners to build out their Stripe capabilities and solutions to grow their business and drive non-linear growth Develop and execute enablement plans Be at the forefront of technical discussions and solution deep-dives with technology partners on the value Stripe can provide to their clients Engage with CTOs, engineering, and other technical leads at ke

S
📍 United States· Full-time
✓ Quality checkedCompany trend -94.6%

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone’s reach while doing the most important work of your career. About the team The Capital team is responsible for managing the end-to-end risk strategy for Stripe's lending product. The team is based in the US and Canada, and is composed of curious, driven, and analytical individuals who are passionate about using their skills to shape the future of Stripe. We partner closely with Stripe's engineering, data science, product, and servicing teams to leverage existing platforms, integrate industry best practices, and develop novel solutions to evaluate and manage credit risk. If you are interested in joining a fast-growing organization and applying your experience to shape the future of Stripe, we encourage you to apply. What you’ll do As a key member of the Capital team, you will have the opportunity to shape the future of Stripe's credit policy by driving meaningful changes to the risk and underwriting framework. You will leverage Stripe's vast data assets to formulate your recommendations and work closely with our partners and, where applicable, leverage third-party data. We are a small and lean team, which means you will have the autonomy and the responsibility to manage end-to-end risk initiatives. Responsibilities Architect and implement credit policies based on Stripe's proprietary data and selected industry data to target, price, and size capital products. Utilize your analytical and technical skills to provide credit risk recommendations, deliver insights, and support strategic business decisions. Collabor

PythonSQLRestMachine Learning
O
📍 United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models. The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient. As the demand for frontier AI grows, so does the complexity of the systems required to support it. Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery. Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows. We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted. About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure. Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online. You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale. This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production. This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key Responsibilities Help bui

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. Working alongside our cloud partners, infrastructure providers, and internal engineering teams, we operate hyperscale AI campuses that support the training and deployment of frontier AI models. The Site Operations team serves as OpenAI's on-site operational presence, helping ensure campuses operate safely, efficiently, and in alignment with Industrial Compute standards. We work closely with Hardware Operations, Infrastructure Delivery, Network Operations, Security, Facilities, Construction, and our infrastructure partners to support day-to-day site execution and maintain operational readiness. As Industrial Compute continues to expand globally, Site Operations plays a critical role in ensuring each campus is prepared to support reliable AI infrastructure at scale. About the Role We are seeking a Site Operations Technician to support the daily operation of Industrial Compute campuses. This role acts as OpenAI's on-site technical representative, helping coordinate activities across hardware operations, facilities, construction, logistics, security, and external service providers. You will perform routine site inspections, support asset tracking, coordinate vendor activities, assist with operational readiness, document site conditions, and help ensure infrastructure issues are identified and resolved quickly. The ideal candidate enjoys working in highly technical environments, is detail-oriented, and thrives in fast-paced operational settings where no two days are the same. Key Responsibilities Perform routine walkthroughs of Industrial Compute facilities to verify operational readiness and identify potential issues. Monitor site conditions and report abnormalities involving hardware spaces, network rooms, utilities, logistics areas, and common infrastructure. Support coordination of vendors, contractors, and partner organizations performing work o

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing

PythonAWSKubernetesRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -83.9%

From $230K/yr

Quick readStrong listing-quality and freshness signals

About the Role The Engineering Acceleration Delivery / Continuous Deployment team builds and operates the systems that safely ship OpenAI’s infrastructure and product code to production. We own the deployment platform, release pipelines, and rollout safety mechanisms that allow engineers across OpenAI to deploy changes rapidly while minimizing operational risk. Our mission is to make production deployments fast, safe, and increasingly autonomous. This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale infrastructure orchestration. In This Role, You Will Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions. Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback. Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows. Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale. Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns. Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks. Build systems that support agent-assisted or autonomous deployment workflows using modern AI tooling. Technologies commonly used in this environment include: Kubernetes for large-scale container orchestration and runtime infrastructure Python and FastAPI for internal services Terraform for infrastructure as code GitOps-based deployment workflows (e.g., ArgoCD, Flux, or similar systems) Buildkite for CI orchestration You may be a strong fit if you: Have worked with Kubernetes-based deployment systems at scale Have experience building or operating continuous deployment platforms Are familiar with GitOps tooling such as

PythonAWSKubernetesGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Stargate and 3P Engineering teams are responsible for building and scaling the external infrastructure ecosystem that powers advanced AI systems. We work across hyperscalers, colocation providers, cloud partners, and strategic third-party operators to turn contracted capacity into production-ready compute. Our scope spans the full lifecycle of external deployments: commercial alignment, technical readiness, network integration, hardware enablement, operational readiness, and long-range scaling strategy. As OpenAI’s infrastructure footprint expands globally, we need leaders who can convert complex partner environments into reliable, high-velocity capacity for training and inference workloads. About the Role We are seeking a Technical Program Manager, Token-as-a-Service (TaaS) to lead delivery of external compute capacity that directly serves OpenAI model workloads. In this role, you will own complex cross-functional programs that transform third-party infrastructure into usable tokens at scale. You will partner across engineering, capacity planning, networking, hardware, finance, product, and external providers to ensure that deployed capacity translates into real production throughput. This role sits at the intersection of infrastructure execution, systems readiness, and business impact. Success requires strong technical fluency, elite program management, and the ability to drive accountability across internal teams and external partners. This is a high-visibility role with direct impact on OpenAI’s ability to scale model training and inference globally. This role is based in San Francisco, CA, with a hybrid work model of 3 days in office per week. Relocation assistance is available. Key Responsibilities Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply. Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments. Build

AWSRestAIRust
S
📍 Bellevue, Washington, United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G

KubernetesAIGoRust
G
📍 New York, NY, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf

🔔

Get new infrastructure sourcing operations lead jobs in United States by email

Daily job updates · Unsubscribe anytime