About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali
Jobs in United States
Infrastructure Sourcing Operations Lead in United States
1,486 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure sourcing operations lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions optimized for advanced AI workloads. We collaborate across research, software, and external hardware partners to design and deploy next-generation AI systems at scale. Our team works closely with silicon vendors and system partners to evaluate emerging technologies, validate performance characteristics, and ensure that hardware capabilities translate effectively to real-world AI workloads. About the Role We are seeking a 3P Hardware Architecture Expert with deep expertise in GPU and accelerator architectures to engage directly with silicon vendors and guide hardware decisions for AI infrastructure. In this role, you will evaluate architectural tradeoffs across compute, memory, and interconnect systems, translating vendor specifications into real-world workload impact. You will play a critical role in early silicon evaluation, benchmarking, and performance validation, helping ensure that next-generation hardware meets the needs of our workloads. This role is highly hands-on and requires both deep technical understanding and the ability to engage at a high level with partners such as NVIDIA and AMD on architectural direction and design tradeoffs. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Engage deeply with silicon vendors (e.g NVIDIA & AMD) on GPU and accelerator architecture tradeoffs. Analyze and interpret performance, power, and efficiency characteristics of next-generation hardware. Translate vendor specifications into expected real-world performance for AI workloads. Evaluate architectural aspects including: compute throughput and utilization memory systems (HBM, cache hierarchies, bandwidth constraints) data types and precision tradeoffs (FP16, BF16, FP8, etc.) interconnect and scaling behavior. Run benchmarks and profiling to validate hardware performance a
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with research, software, and external hardware partners to shape the next generation of AI systems, from silicon through full-scale deployments. Our team focuses on understanding and optimizing performance across the full system stack—ensuring that architectural decisions are grounded in rigorous, quantitative analysis of real-world workloads. About the Role We are seeking a Performance Modeling Lead to build and lead a small, high-impact team responsible for answering forward-looking architectural questions across AI infrastructure systems. You will develop modeling frameworks and methodologies to evaluate system-level tradeoffs and guide key design decisions. Your work will directly influence reference architectures, vendor designs, and long-term infrastructure strategy. This role sits at the intersection of AI workloads, system architecture, and quantitative modeling, and requires strong technical judgment, ownership, and the ability to translate complex analysis into clear, actionable guidance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Build and own a performance modeling framework/toolchain to evaluate AI systems across multiple levels of abstraction. Analyze and quantify architectural tradeoffs across compute, memory, networking, storage, and system topology. Develop performance models to guide decisions on: scale-up vs. scale-out architectures interconnect and network design memory hierarchy and system balance. Translate modeling outputs into clear recommendations for internal teams and external hardware vendors. Influence reference designs and vendor roadmaps through data-driven insights. Partner closely with machine learning, systems, and hardware teams to understand workload characte
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation
NVIDIA silicon runs the world's AI infrastructure. The frequency it delivers across every voltage, process corner, and workload is not assumed. It is measured, correlated, and validated. This role does that work. The Silicon Co-Design Group is where architecture intent becomes silicon reality. We own the boundary between what was designed and what was built, and we are the team that knows the difference. When a program ships at frequency and at quality, this team is a reason why. You will be the person who follows through between simulation and silicon. When the model is wrong, a frequency corner that doesn't hold, a Vmin that walks, a critical path that timing analysis missed, you find out why, and your data is what the rest of the program acts on. Architecture, design, and product teams do not guess. They use your numbers. The engineers who do this well are rare. They think like circuit designers, work like experimentalists, and reason like data scientists. If that is you, read on. What you'll be doing: Own silicon speed characterization from first power-on through production sign-off, covering frequency, Vmin, Vmax, and timing margins across the full PVT space. Close the correlation gap. Tie pre-silicon timing analysis and critical path predictions to measured silicon, quantify where the model diverges from reality, and produce analysis that architecture and design can act on with confidence. Trace failures to their source, whether a microarchitectural bottleneck, a critical path that doesn't close under voltage, a clocking issue, or a process corner the model didn't anticipate, and drive the resolution. Own and build AI agents that work at your direction: automated test orchestration, intelligent data pipelines, and analysis flows that expand coverage and compress cycle time without sacrificing difficulty. Know where AI accelerates real work and whe
$99.2K – $165.4K/yr
Role Summary The CISO, Infrastructure and Cloud Services organization enables Pfizer’s mission by delivering secure, resilient, and scalable technology platforms that support operations and future growth. Through a unified operating model, it strengthens accountability, accelerates decision‑making, and improves availability, security, and cost efficiency. The CISO Marketing Manager will support the Senior Manager, Marketing and Enterprise Awareness to develop meaningful marketing materials that tell the CISO story to external clients by executing strategic, creative, and visually compelling communications including presentations, graphics, pictures, charts and tables that support cybersecurity awareness, organizational priorities, and business outcomes. The manager will lead the creation of a marketing strategy for the Chief Information Security Office Team. Role Responsibilities Lead the development and execution of integrated marketing and communications strategies in support of CISO priorities, strategic initiatives, and portfolio management objectives. Produce high-quality written, visual, and digital content that strengthens awareness, engagement, and understanding of cybersecurity programs across the organization. Create and manage internal communications, including newsletters, leadership messages, campaign materials, SharePoint content, presentations, infographics, and other digital assets. Simplify complex cybersecurity and technical information into clear, engaging, and actionable content for diverse audiences, including executives, business partners, and internal colleagues. Partner closely with cross-functional teams across Cyber Defense, Infrastructure, Cloud, and related Digital and Technology groups to ensure consistent, coordinated, and impactful messaging. Develop executive-ready presentations and visual storytelling materials
About the Role As Head of Finance - Data Centers, you will be the finance leader for OpenAI’s self-built data center efforts. You will partner with teams across infrastructure, real estate, energy, construction, procurement, and finance to turn proposed sites into sound investment decisions and funded projects. This role spans the full development lifecycle: evaluating opportunities, building investment cases, forecasting capital needs, managing construction budgets, and helping determine how projects should be financed. You will give leadership a clear view of project economics, funding requirements, and risks as we build data center capacity at scale. In this role, you will: Lead financial evaluation of proposed data center and related power infrastructure projects, including site economics, development costs, capacity phasing, lifecycle costs, and key risks. Build and own project-level models and capital expenditure forecasts that connect construction schedules, power delivery, equipment procurement, contingencies, and funding needs. Establish capital budgets and financial controls for active builds. Track commitments, actual spending, change orders, and forecasts to completion; identify cost or schedule risks early. Partner with development, engineering, energy, construction, and procurement leaders on decisions that affect cost, timing, and long-term performance. Work with Treasury, Corporate Finance, Tax, and Legal to evaluate financing options, including project or construction debt, leases, joint ventures, and other partnership structures where appropriate. Prepare investment recommendations and capital approval materials for senior leadership, translating complex project details into clear choices and tradeoffs. Build a consistent portfolio view of project costs, cash requirements, milestones, and financial performance. Partner with Accounting and operations teams through project completion and handoff. You might thrive in this role if you have: Prior exper
From $131K/yr
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models. We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience. You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments. Key Responsibilities Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency. Develop forecasting models for inference demand across products, regions, and model families. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs. Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions. Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps. Communicate technical findings clearly to both engineering teams and executive leadership. Qualifications MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience). 5+ years of experience working in the infrastructure data science space. Strong ex
About the Team OpenAI is building the world’s most advanced AI infrastructure ecosystem. The Site Readiness & Development team owns the upstream diligence and development work required to convert powered-land opportunities into executable infrastructure options. About the Role The Site Selection Lead sets the site-selection strategy for powered-land and greenfield/brownfield opportunities. This role defines, with cross-functional partners, what makes a site attractive, executable, and scalable, then turns that shared rubric into portfolio choices and deployment plans. Key Responsibilities Own the site-selection strategy across priority geographies, including market maps, pipeline segmentation, source refreshes, live-location tracking, and portfolio prioritization. Convene Power, Land, Development, Engineering, Construction, Environmental, Community, Commercial, Legal, and Finance partners to define the shared rubric for a high-quality site. Translate that rubric into clear screening criteria across power readiness, land control, timing, scale, cost, community risk, environmental constraints, AHJ path, budget, schedule, and counterparty credibility. Orchestrate Commercial, Legal, Finance, Power, and Development teams to shape exclusivity, initial terms, site-control strategy, funding assumptions, and the path to deployment before deeper commitment. Build a deployment plan for top candidates that identifies capacity, sequencing, capital implications, critical-path approvals, risks, owners, and stage-gate decisions. Lead conceptual site planning and test-fit screening using boundaries, map layers, buildings, substations, roads, parking, laydown, stormwater, utility corridors, acreage, and estimated load. Maintain a single decision-support view of pipeline stage, site health, power readiness, finance review, maps, project evidence, outstanding source data, owners, deadlines, and escalations. Present executive-ready advance, defer, or stop recommendations with clear
Who we are About Stripe Stripe is a technology company that builds economic infrastructure for the internet. Businesses of all sizes—from new startups to public companies like Amazon and Salesforce—use our software to accept online payments and run technically sophisticated financial operations in more than 100 countries. By helping businesses accept payments from anywhere in the world and accelerating the economic shift from offline to online, we aim to increase the GDP of the internet. About the team Stripe’s Leadership Recruiting Team (LRT) is a high-performing group of strategic and creative problem solvers. We are passionate about finding and hiring the very best leaders and experts from around the globe to help drive the next arc of Stripe’s growth. What you’ll do In this role, you will be directly responsible for partnering with senior stakeholders to identify, attract, assess, and hire individuals for leadership roles across the company. Responsibilities Strategically partner with executives to hire top talent for Stripe’s highest priority engineering leadership roles Define, design, and implement recruiting processes and search engagements for a variety of senior leadership roles. This includes understanding and mapping the talent landscape, internal calibration and referrals, and matching those data points with the external talent supply Manage end-to-end leadership-level searches from inception of candidate contact to onboarding, including research, outreach strategies, candidate experience, negotiations, partnering with HR for successful onboarding, and communicating with stakeholders along the way Collaborate with cross-functional partners, such as People Partners, Compensation, and Business Leads, to align hiring goals with business needs Develop competency-based interviews that surface a candidate’s full range of capabilities, explore their aspirations, and gain insight into how their experience will apply to a leadership role at Stripe Fo
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b
NVIDIA is seeking an a PCB Library Engineer to join our PCB Design Infrastructure team. In this role, you will help develop and maintain the PCB library assets used across NVIDIA's Data Center, AI, Networking, Automotive, and Graphics products. Working alongside experienced PCB designers, library engineers, mechanical engineers, manufacturing engineers, and component engineers, you will create and validate component footprints, schematic symbols, mechanical components, panel definitions, and other critical design assets that enable successful product development. This position provides an excellent opportunity to build expertise in PCB design, manufacturing, component engineering, and design automation while supporting some of the most advanced computing platforms in the world. What you'll be doing: Develop PCB footprints, padstacks, schematic symbols, and mechanical library content using Cadence PCB design tools. Review component datasheets, package drawings, and engineering specifications to create accurate design libraries. Support library verification, release, and documentation processes. Partner with PCB design, mechanical engineering, manufacturing engineering, and operations teams to resolve library-related issues. Learn and apply industry standards, including IPC requirements, DFM, DFA, and DFT principles. Support quality initiatives to ensure library content is accurate, manufacturable, and scalable. Participate in continuous improvement and automation efforts within the library environment. Develop a strong understanding of PCB fabrication, assembly, and component technologies. What We Need to See: BS degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Manufacturing Engineering, or a related field or equivalent experience. <p
Other cities to consider
More places hiring for this role
Get new infrastructure sourcing operations lead jobs in United States by email
Daily job updates · Unsubscribe anytime