About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and inform next-generation infrastructure design. About the Role We are seeking Performance Modeling Engineers to develop and apply modeling tools that evaluate AI system performance and inform architectural decisions. In this role, you will work closely with the Performance Modeling Lead and partner teams to analyze system behavior, run simulations or analytical models, and help quantify tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks and applying them to real-world questions that impact system design and vendor decisions. This role is well-suited for engineers with strong software or modeling backgrounds who are interested in developing deeper expertise in system architecture and AI infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Develop and maintain performance modeling tools and frameworks. Build models to evaluate system behavior across: compute, memory, and interconnect subsystems distributed system scaling and bottlenecks. Run simulations and analytical models to support architectural tradeoff analysis. Collaborate with performance modeling lead and system architects to answer forward-looking design questions. Analyze and interpret modeling outputs, translating results into actionable insights. Validate models against real system measurements and workload behavior. Contribute to improving modeling fidelity, usability, and scalability. Qualifications Strong software engineeri
Jobs in United States
Infrastructure Security Engineer in United States
1,531 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure security engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with research, software, and external hardware partners to shape the next generation of AI systems, from silicon through full-scale deployments. Our team focuses on understanding and optimizing performance across the full system stack—ensuring that architectural decisions are grounded in rigorous, quantitative analysis of real-world workloads. About the Role We are seeking a Performance Modeling Lead to build and lead a small, high-impact team responsible for answering forward-looking architectural questions across AI infrastructure systems. You will develop modeling frameworks and methodologies to evaluate system-level tradeoffs and guide key design decisions. Your work will directly influence reference architectures, vendor designs, and long-term infrastructure strategy. This role sits at the intersection of AI workloads, system architecture, and quantitative modeling, and requires strong technical judgment, ownership, and the ability to translate complex analysis into clear, actionable guidance. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Build and own a performance modeling framework/toolchain to evaluate AI systems across multiple levels of abstraction. Analyze and quantify architectural tradeoffs across compute, memory, networking, storage, and system topology. Develop performance models to guide decisions on: scale-up vs. scale-out architectures interconnect and network design memory hierarchy and system balance. Translate modeling outputs into clear recommendations for internal teams and external hardware vendors. Influence reference designs and vendor roadmaps through data-driven insights. Partner closely with machine learning, systems, and hardware teams to understand workload characte
NVIDIA silicon runs the world's AI infrastructure. The frequency it delivers across every voltage, process corner, and workload is not assumed. It is measured, correlated, and validated. This role does that work. The Silicon Co-Design Group is where architecture intent becomes silicon reality. We own the boundary between what was designed and what was built, and we are the team that knows the difference. When a program ships at frequency and at quality, this team is a reason why. You will be the person who follows through between simulation and silicon. When the model is wrong, a frequency corner that doesn't hold, a Vmin that walks, a critical path that timing analysis missed, you find out why, and your data is what the rest of the program acts on. Architecture, design, and product teams do not guess. They use your numbers. The engineers who do this well are rare. They think like circuit designers, work like experimentalists, and reason like data scientists. If that is you, read on. What you'll be doing: Own silicon speed characterization from first power-on through production sign-off, covering frequency, Vmin, Vmax, and timing margins across the full PVT space. Close the correlation gap. Tie pre-silicon timing analysis and critical path predictions to measured silicon, quantify where the model diverges from reality, and produce analysis that architecture and design can act on with confidence. Trace failures to their source, whether a microarchitectural bottleneck, a critical path that doesn't close under voltage, a clocking issue, or a process corner the model didn't anticipate, and drive the resolution. Own and build AI agents that work at your direction: automated test orchestration, intelligent data pipelines, and analysis flows that expand coverage and compress cycle time without sacrificing difficulty. Know where AI accelerates real work and whe
About the Role As Head of Finance - Data Centers, you will be the finance leader for OpenAI’s self-built data center efforts. You will partner with teams across infrastructure, real estate, energy, construction, procurement, and finance to turn proposed sites into sound investment decisions and funded projects. This role spans the full development lifecycle: evaluating opportunities, building investment cases, forecasting capital needs, managing construction budgets, and helping determine how projects should be financed. You will give leadership a clear view of project economics, funding requirements, and risks as we build data center capacity at scale. In this role, you will: Lead financial evaluation of proposed data center and related power infrastructure projects, including site economics, development costs, capacity phasing, lifecycle costs, and key risks. Build and own project-level models and capital expenditure forecasts that connect construction schedules, power delivery, equipment procurement, contingencies, and funding needs. Establish capital budgets and financial controls for active builds. Track commitments, actual spending, change orders, and forecasts to completion; identify cost or schedule risks early. Partner with development, engineering, energy, construction, and procurement leaders on decisions that affect cost, timing, and long-term performance. Work with Treasury, Corporate Finance, Tax, and Legal to evaluate financing options, including project or construction debt, leases, joint ventures, and other partnership structures where appropriate. Prepare investment recommendations and capital approval materials for senior leadership, translating complex project details into clear choices and tradeoffs. Build a consistent portfolio view of project costs, cash requirements, milestones, and financial performance. Partner with Accounting and operations teams through project completion and handoff. You might thrive in this role if you have: Prior exper
About the Team Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior. About the Role You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong. This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware. We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack. In this role, you will: Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows. Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers. Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware,
About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models. We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience. You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments. Key Responsibilities Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency. Develop forecasting models for inference demand across products, regions, and model families. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs. Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions. Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps. Communicate technical findings clearly to both engineering teams and executive leadership. Qualifications MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience). 5+ years of experience working in the infrastructure data science space. Strong ex
About the Team OpenAI is building the world’s most advanced AI infrastructure ecosystem. The Site Readiness & Development team owns the upstream diligence and development work required to convert powered-land opportunities into executable infrastructure options. About the Role The Site Selection Lead sets the site-selection strategy for powered-land and greenfield/brownfield opportunities. This role defines, with cross-functional partners, what makes a site attractive, executable, and scalable, then turns that shared rubric into portfolio choices and deployment plans. Key Responsibilities Own the site-selection strategy across priority geographies, including market maps, pipeline segmentation, source refreshes, live-location tracking, and portfolio prioritization. Convene Power, Land, Development, Engineering, Construction, Environmental, Community, Commercial, Legal, and Finance partners to define the shared rubric for a high-quality site. Translate that rubric into clear screening criteria across power readiness, land control, timing, scale, cost, community risk, environmental constraints, AHJ path, budget, schedule, and counterparty credibility. Orchestrate Commercial, Legal, Finance, Power, and Development teams to shape exclusivity, initial terms, site-control strategy, funding assumptions, and the path to deployment before deeper commitment. Build a deployment plan for top candidates that identifies capacity, sequencing, capital implications, critical-path approvals, risks, owners, and stage-gate decisions. Lead conceptual site planning and test-fit screening using boundaries, map layers, buildings, substations, roads, parking, laydown, stormwater, utility corridors, acreage, and estimated load. Maintain a single decision-support view of pipeline stage, site health, power readiness, finance review, maps, project evidence, outstanding source data, owners, deadlines, and escalations. Present executive-ready advance, defer, or stop recommendations with clear
Who we are About Stripe Stripe is a technology company that builds economic infrastructure for the internet. Businesses of all sizes—from new startups to public companies like Amazon and Salesforce—use our software to accept online payments and run technically sophisticated financial operations in more than 100 countries. By helping businesses accept payments from anywhere in the world and accelerating the economic shift from offline to online, we aim to increase the GDP of the internet. About the team Stripe’s Leadership Recruiting Team (LRT) is a high-performing group of strategic and creative problem solvers. We are passionate about finding and hiring the very best leaders and experts from around the globe to help drive the next arc of Stripe’s growth. What you’ll do In this role, you will be directly responsible for partnering with senior stakeholders to identify, attract, assess, and hire individuals for leadership roles across the company. Responsibilities Strategically partner with executives to hire top talent for Stripe’s highest priority engineering leadership roles Define, design, and implement recruiting processes and search engagements for a variety of senior leadership roles. This includes understanding and mapping the talent landscape, internal calibration and referrals, and matching those data points with the external talent supply Manage end-to-end leadership-level searches from inception of candidate contact to onboarding, including research, outreach strategies, candidate experience, negotiations, partnering with HR for successful onboarding, and communicating with stakeholders along the way Collaborate with cross-functional partners, such as People Partners, Compensation, and Business Leads, to align hiring goals with business needs Develop competency-based interviews that surface a candidate’s full range of capabilities, explore their aspirations, and gain insight into how their experience will apply to a leadership role at Stripe Fo
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence. About the Role We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly. Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In This Role, You Will Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry b
NVIDIA is seeking an a PCB Library Engineer to join our PCB Design Infrastructure team. In this role, you will help develop and maintain the PCB library assets used across NVIDIA's Data Center, AI, Networking, Automotive, and Graphics products. Working alongside experienced PCB designers, library engineers, mechanical engineers, manufacturing engineers, and component engineers, you will create and validate component footprints, schematic symbols, mechanical components, panel definitions, and other critical design assets that enable successful product development. This position provides an excellent opportunity to build expertise in PCB design, manufacturing, component engineering, and design automation while supporting some of the most advanced computing platforms in the world. What you'll be doing: Develop PCB footprints, padstacks, schematic symbols, and mechanical library content using Cadence PCB design tools. Review component datasheets, package drawings, and engineering specifications to create accurate design libraries. Support library verification, release, and documentation processes. Partner with PCB design, mechanical engineering, manufacturing engineering, and operations teams to resolve library-related issues. Learn and apply industry standards, including IPC requirements, DFM, DFA, and DFT principles. Support quality initiatives to ensure library content is accurate, manufacturable, and scalable. Participate in continuous improvement and automation efforts within the library environment. Develop a strong understanding of PCB fabrication, assembly, and component technologies. What We Need to See: BS degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, Manufacturing Engineering, or a related field or equivalent experience. <p
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
About the Team Industrial Compute is building the world's most advanced AI infrastructure. Working alongside our capital partners, engineering teams, and construction organizations, we design and deliver large-scale, mission-critical compute campuses that power the next generation of AI. Our Design organization brings together engineering, construction, and digital design to ensure facilities are coordinated, constructible, and optimized from the earliest planning phases through deployment. About the Role We are seeking a BIM Designer & Coordinator to support the planning and design of large-scale industrial and mission-critical facilities. In this role, you will develop, coordinate, and maintain multidisciplinary BIM models across Civil, Electrical, Mechanical, Architectural, and Structural disciplines, enabling early design validation, constructability reviews, equipment planning, and cross-functional coordination. You will partner closely with internal engineering teams, external design consultants, contractors, and project stakeholders to produce coordinated BIM deliverables that improve design quality, reduce project risk, and support efficient execution across Industrial Compute's global infrastructure portfolio. Key Responsibilities: Develop and maintain conceptual and schematic BIM models for large-scale industrial and MEP-intensive facilities. Coordinate BIM models across Civil, Electrical, Mechanical, Architectural, and Structural disciplines. Create and maintain federated models used for design reviews, spatial coordination, constructability analysis, and clash detection. Model major building systems including equipment layouts, utility corridors, electrical rooms, mechanical rooms, structural framing, site infrastructure, and architectural constraints. Coordinate equipment clearances, maintenance access, routing zones, shafts, risers, utility entrances, and major MEP pathways. Translate engineering sketches, basis-of-design documents, equipment lists
About the Team OpenAI’s Finance and Revenue Operations organization builds the commercial infrastructure that enables the business to scale with speed, discipline, and financial integrity. Within Revenue Operations, Deal Desk partners closely with Sales, Partnerships, Product, Engineering, Legal, Technical Revenue, Finance, Billing, Order Management, and GTM Systems to turn complex commercial opportunities into executable, scalable transactions. About the Role We are hiring a Strategic Deals & Commercial Architecture Lead — Marketplaces & Partnerships to own the commercial architecture, execution, and governance layer for marketplace-enabled and partner-led transactions. This senior individual-contributor role is for someone who combines enterprise deal judgment, marketplace fluency, analytical rigor, and a builder mindset. You will take shaped opportunities from intake through approval and launch readiness, translating first-of-kind structures into clear economics, executable terms, quote-to-cash requirements, controls, and operating mechanisms. You will partner closely with teams that own business development and partner relationships; this role does not own partner sourcing, pipeline generation, or sales closing. You will support the commercial requirements for onboarding and integration without owning technical delivery. Success is measured both by the decisions and deals you enable and by the durable policies, systems, and controls you leave behind. This role is based in San Francisco, CA. We use a hybrid work model of three office days per week and offer relocation assistance. In This Role, You Will Lead the commercial architecture and execution of complex marketplace-enabled and partner-led enterprise transactions from intake through approval, contracting, launch readiness, and operational handoff. Structure private offers, pricing, fees, incentives, commitments, revenue share, credits, renewals, amendments, and other non-standard or multi-party terms
From $126.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Data Center Engineer , you'll help us scale our Core/Edge Data Centers and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Technical Lead Data Center Engineer. You will: Develop and maintain the Core/Edge Data Center and hardware infrastructure to meet the large scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental life cycles. Own efforts to track and mitigate systemic issues preventing hosts from returning to service. Identify and solve critical problems and prevent them from re-occurring via root cause analysis and giving rec
We’re on a mission to build the world’s first AI Development Infrastructure: software that enables enterprises to safely evolve from today’s AI assistants to tomorrow’s autonomous workflows, all while maintaining full visibility and control. We’re creating a new category, and we need a Product Marketing leader to help define it. You’ll shape and scale Coder’s story across every touchpoint, building thought leadership, deepening our enterprise presence, and influencing how the world thinks about AI development. You’ll define the vision and strategy for the team, working across the company to drive clarity, alignment, and growth, and position Product Marketing as a force multiplier for our business. This role is perfect for a player-coach: someone who can translate complex technology into clear, credible stories, lead insightful analyst conversations, and dive deep into customer segmentation - all in a day’s work. What you’ll do here Evaluate, refine, and validate Coder’s messaging and positioning to explain how we fit into enterprise AI transformations. Build clear, visual frameworks that communicate complex design choices and product architecture to technical decision-makers. Drive alignment and consistency across internal teams and customer touchpoints through continuous iteration. Partner with Sales, Customer Solutions, and Product to identify and operationalize emerging use cases for Coder at the frontier of agentic development. Grow the Product Marketing team while providing strategic guidance on scaling thought leadership, developer relations, and enablement to create cohesive marketing to our ICP. Drive a unified technical content strategy with Marketing, Product, and Customer Solutions that connects product truth to customer needs and market opportunities, strengthening Coder’s technical credibility. What we’re looking for 10+ years of relevant experience, with 3+ years in a leadership/management role Action‑oriented; execute autonomously in fast-changing env
Other cities to consider
More places hiring for this role
Get new infrastructure security engineer jobs in United States by email
Daily job updates · Unsubscribe anytime