Jobs in United States

Infrastructure Engineer in United States

1,504 active opportunities · Updated October 2026

Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

M
📍 United States· Full-time
✓ High-confidence listingCompany trend -97.2%

From $126K/yr

Quick readStrong listing-quality and freshness signals

Senior Forward Deployed Engineers (FDE) sit at the intersection of enterprise customer environments and a fast-moving internal product. They partner directly with customers to design, build, troubleshoot, and improve production solutions, and they are ultimately accountable for helping customers get to production. This role is best suited for engineers who want to own technical outcomes end to end: understanding what a customer is trying to build, shipping the integration, and feeding learnings back into the product roadmap. This role will be based remotely in the United States (East Coast). Responsibilities Customer success Serve as the primary technical owner for customer engagements from initial discovery through production rollout Understand each customer's architecture, constraints, and definition of success, and drive toward that outcome Manage expectations, communicate risks clearly, and help customers navigate technical decisions with confidence Technical integration Build the connectors, pipelines, and supporting tooling needed to make the platform work inside real enterprise environments Write production-quality code, troubleshoot issues, and implement fixes directly in active workstreams Work effectively within customer environments that have different stacks, infrastructure, and integration constraints Product feedback loop Capture product feedback with precision, including logs, reproduction steps, and a clear proposed path forward Use customer engagements to identify product gaps, surface recurring patterns, and help improve the product roadmap Document technical decisions and tradeoffs clearly so product and engineering teams can extend the work Workstream collaboration Partner closely with product and engineering teams to help turn field patterns into reusable product capabilities Contribute directly in focused workstreams by helping drive technical design, implementation, and delivery Make sound engine

MongoDBAWSAzureAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Enterprise Identity team builds the identity foundation that enables organizations to adopt and use OpenAI products securely and reliably. The team owns the enterprise identity stack, including SSO, SCIM, tenant architecture, and identity capabilities across the enterprise admin experience and OpenAI's growing multi-product portfolio. About the Role We are looking for a hands-on senior technical leader to own the architecture and evolution of OpenAI's Enterprise Identity systems. You will set the long-term technical vision for the entire stack, establish shared identity primitives across products, and be accountable for systems that are foundational to our enterprise business. This role requires operating well beyond a single service or feature area. You will identify the most consequential architectural investments, align teams around durable solutions, and ensure our identity platform meets an exceptionally high bar for scale, availability, latency, and security. This role will be based in our San Francisco or Mountain View office. In this role, you will: Own the technical vision and architecture for the Enterprise Identity stack, including SSO, SCIM, tenant architecture, groups, permissions, and identity capabilities in enterprise administration surfaces. Lead the design and evolution of highly available, latency-sensitive identity systems serving a large and diverse global enterprise customer base. Establish common identity models and primitives that work consistently across OpenAI's products and enable the organization to scale. Set a high security bar by anticipating abuse cases, failure modes, and the long-term implications of new capabilities. Drive alignment across enterprise product, infrastructure, and security partners, resolving ambiguity and influencing roadmaps beyond the immediate team. Provide technical leadership to senior engineers and raise the quality of architecture and execution across the broader organization. You might thr

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Legal team is building the next generation of AI-powered products and experiences for the legal industry. We are exploring how advanced AI systems can transform legal workflows, improve access to information, and enable legal professionals and organizations to work more effectively. As a founding member of the Legal engineering team, you will help define the technical foundation for this new product area from the earliest stages. You’ll operate at the intersection of AI, product, and real-world legal workflows—identifying opportunities, building prototypes, and turning emerging ideas into scalable products that can create meaningful impact. We operate with a startup-like mindset inside OpenAI: small teams, rapid iteration cycles, and a willingness to explore bold ideas, learn quickly, and adapt based on user feedback. Our goal is to build products that meaningfully improve how legal professionals work while leveraging OpenAI’s cutting-edge models and infrastructure. About the Role As a Founding Full-Stack Software Engineer on the Legal team, you will help imagine, build, and scale new AI-powered products for the legal industry. You’ll work across the stack to design intuitive user experiences, build robust backend systems, and create the foundations for products used by legal professionals and organizations around the world. You’ll have significant ownership from the earliest stages—working closely with product, design, research, and go-to-market partners to understand customer needs, shape product direction, and deliver high-impact solutions. This includes rapidly prototyping new concepts, building production-quality applications on top of OpenAI’s platforms, and developing new technical approaches when existing systems are not sufficient. We’re looking for engineers who thrive in ambiguity, have strong product instincts, and enjoy building from 0→1. You should be comfortable moving quickly, making thoughtful technical decisions, and taking owner

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI’s rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI’s third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure. Key Responsibilities Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. Develop operationa

AWSAzureGCPRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for an experienced systems software engineer to help define and build the host software stack for our custom next-generation AI systems. You will work close to the hardware on performance-critical software, including Linux kernel drivers, high-throughput I/O paths, and system-scale networking and RDMA. This role spans architecture, implementation, platform bring-up, debugging, and performance optimization. You will work across hardware and software boundaries to make new systems usable end to end, from low-level device interfaces through userspace tooling and production validation. In this role you will: Design, implement, and debug host-side systems software for AI infrastructure, including Linux kernel drivers and supporting userspace components. Build and optimize software paths for high-throughput, low-latency communication, including RDMA and related networking functionality. Develop software around PCIe, DMA, NICs, accelerators, memory movement, and device interaction. Bring up new hardware platforms and diagnose complex issues across kernel, firmware, networking, and hardware boundaries. Build tooling for integration, testing, diagnostics, observability, qualification, and performance characterization. Collaborate with hardware, networking, and platform teams to define interfaces and integrate new capabilities. Work with external vendors where needed to integrate technologies and drive issues to resolution. Contribute across the systems sof

PythonAWSLinuxRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects, and spend less time on repetitive tasks. AI has already transformed how code is written, but software engineering extends far beyond coding. Our mission is to apply AI across the entire software development lifecycle (SDLC) — from design and implementation to code review, testing, debugging, issue remediation, maintenance, documentation, and user support. The team is also responsible for developer-facing Codex experiences including the Codex IDE Extension and the terminal interface, which are used daily by developers ranging from individual open-source contributors to some of the world’s largest engineering organizations. The team also works closely with the open-source software community, building tools that help maintainers and contributors manage increasingly complex projects. We believe AI can make open-source development more sustainable by reducing the operational burden of reviewing contributions, triaging issues, maintaining quality, and supporting growing communities. By building the future of software development, we're helping advance OpenAI's mission of ensuring that the benefits of AI reach people around the world. About the Role We’re hiring a Full Stack Software Engineer to help invent the next generation of AI-powered software development workflows. “Full stack” in this role means much more than traditional frontend and backend development. You'll own complete product experiences, spanning user interfaces, workflow orchestration, agent and prompt design, backend systems, and cloud infrastructure. This is a highly product-oriented role. You'll work directly on the workflows developers use every day, identifying bottlenecks and rethinking how software gets built in a world where AI agents are active participants in the development process. The features you ship will inf

TypeScriptAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a

AWSRestAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

Our work at NVIDIA is dedicated towards a computing model focused on visual and AI computing. For two decades, NVIDIA has pioneered visual computing, the art and science of computer graphics, with our invention of the GPU. The GPU has also shown to be spectacularly effective at solving some of the most complex problems in computer science. Today, NVIDIA’s GPU simulates human intelligence, running deep learning algorithms and acting as the brain of computers, robots and self-driving cars that can perceive and understand the world. We are looking to grow our company and teams with the smartest people in the world and there has never been a more exciting time to join our team! The AI Infrastructure Product Design team creates software used by engineers and researchers to prepare data, run complex workflows, and understand results. This internship offers ownership of a defined product problem from early research through a tested design and implementation handoff. Designers on this team often move between Figma and working HTML prototypes, and may hand off HTML directly to engineering. This work calls for a high standard of visual and interaction design alongside technical fluency. The role is a good fit for someone who enjoys making technically complex systems easier to understand and who uses large language models and software agents thoughtfully as part of their design and prototyping process. What you will be doing: Own a focused design project for an internal AI infrastructure product, from understanding the problem through a validated design and implementation handoff. Interview engineers and researchers, map their workflows, and turn the findings into clear product requirements, user flows, and interaction models. Create precise, implementation-ready interface designs and interactive prototypes in Figma and HTML/CSS, with careful attention to typography, hierarchy, spacing, visual consistency, interactio

JavaScriptReactAI
C
📍 Tampa Florida United States, United States
✓ Quality checkedCompany trend +800%

At Citi Services - Global Trade and Working Capital Solutions (TWCS) Technology Organization, we are on a mission to harness the power of data to drive innovation, create exceptional customer experiences, and solve complex business challenges. Our data team is at the heart of this mission, building the scalable and resilient infrastructure that turns data into our most asset. We are a passionate, collaborative group dedicated to pushing the boundaries of what's possible. The Opportunity We are seeking an Engineering Director to join our development team. The ideal candidate is a seasoned technologist with extensive experience in building and delivering data & AI solutions for business functions. This individual will be directly and fully accountable for solution delivery and should have a history of creating strategic technology architecture roadmaps aligned with business outcomes. The successful candidate will define and execute the technology roadmap for the Data & AI portfolio of TWCS, providing strategic direction and critical input into technology decisions to ensure a scalable buildout. This role involves building strong relationships with senior business and technology partners, driving agile execution, and leading a team of expert engineers to deliver with velocity and quality. Applicants should demonstrate exceptional technical acumen, a strong data engineering background, and a proven ability to lead and provide technical direction to high-performing engineering teams. A proven expertise in Data Engineering, Data Analytics, and AI Engineering delivery at scale is essential. Responsibilities: Manage/develop multiple teams of professionals to accomplish established goals and conduct personnel duties for team (e.g. performance evaluations, hiring and disciplinary actions) as well as ensure team adheres to best practices and processes Develop vision for team aroun

PythonJavaAWSAzure
H
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Hyliion is committed to creating innovative solutions that enable clean, flexible and affordable electricity production. The Company’s primary focus is to develop distributed power generators that can operate on various fuel sources to future-proof against an ever-changing energy economy. Job Purpose The Senior Manager, Manufacturing Engineering owns the manufacturing engineering function across Hyliion's Cedar Park, TX headquarters and primary production site and its Milford, OH R&D operations. The role is the bridge between the engineering design of the KARNO Power Module and the scalable, repeatable production system required for commercial ramp, defining and driving the processes, tooling, documentation, and infrastructure that transform KARNO from a highly engineered prototype into a manufacturable product. This is a player-coach role: the incumbent builds and leads a team of manufacturing engineers while remaining personally hands-on in the most challenging production problems. AI at Hyliion At Hyliion, AI is core to how we work. We equip every team member with leading AI tools and count on you to use them — to move faster, solve harder problems, and help us realize the full potential of KARNO technology for the world. Duties and Responsibilities Productionization: translate engineering designs into manufacturable, repeatable, and scalable production configurations across both sites, aligned to the commercial ramp roadmap. Tooling and process development: design, develop, and qualify manufacturing tooling, fixtures, and processes that support quality, efficiency, and volume scalability. Manufacturing documentation: own PFMEAs, control plans, work instructions, SOPs, and MBOMs, keeping them accurate, current, and accessible to the production floor. Lean and material flow: define and implement standard work, material flow, and handling strategies that increase throughput, reduce waste, and prepare the factory for

AIGoExcelSEM
T
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is seeking a SoC Design Verification Engineer to validate the System Management Controller (SMC) and enable seamless multi-chip integration. In this role, you will design and execute tests, build infrastructure, and debug issues across chiplet-based SoCs. You’ll have the opportunity to work with remote mentorship while contributing to the foundation of scalable multi-die systems. This role is hybrid, based out of Toronto, Ontario, Boston, MA or Santa Clara, CA. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Proficient in SystemVerilog, SV-UVM, Python, and C/C++ with strong verification skills. Experienced in writing test plans, building infrastructure, and debugging hardware/software flows. Comfortable working with remote mentorship and distributed teams. Familiar with AI-assisted tools like Copilot, Cursor, and Claude to accelerate verification. What We Need Develop and maintain SMC tests and supporting DV infrastructure. Write, execute, and track test plans for chiplet and multi-chip SoC designs. Use C/C++ to develop tests compiled, loaded, and executed directly on the DUT. Triage, analyze, and debug issues in clos

PythonAWSAIC++
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -83.9%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's Industrial Compute organization builds and operates the infrastructure required to train and serve frontier AI models. The Capacity Planning team connects rapidly changing research and product demand with the compute, networking, storage, power, data center, hardware, and operational resources required to make that demand executable. About the Role We are seeking a Technical Program Manager to build and lead capacity planning across OpenAI's large-scale AI infrastructure. You will translate uncertain workload demand into clear infrastructure requirements, allocation decisions, supply commitments, activation priorities, and long-range capacity strategies. This role sits at the intersection of research, engineering, infrastructure, finance, sourcing, deployment, and operations. You will create the planning models, operating cadences, governance mechanisms, and source-of-truth systems that allow teams to understand what capacity is required, what is available, what is at risk, and what decisions must be made. This is not a finance-only forecasting or reporting role. Success requires technical fluency across the infrastructure stack, strong analytical judgment, and the ability to move consequential decisions forward when requirements, timelines, and supply conditions change quickly. Key Responsibilities Own capacity-planning processes across near-term workload allocation, quarterly execution, and longer-range infrastructure horizons. Translate research, training, inference, and product demand into compute, accelerator, cluster, networking, storage, rack, power, and site requirements. Develop scenarios that make assumptions, confidence levels, constraints, sensitivities, and decision points explicit. Reconcile requested demand against contracted, delivered, installed, activated, and workload-usable capacity. Partner with research and engineering teams to understand workload priorities, technical dependencies, utilization patterns, and changing req

PythonSQLAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -83.9%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences. As frontier models and agents evolve how they use memory, history and snapshots, the underlying architecture increasingly shapes the capabilities products can deliver—and their latency, reliability, cost and efficiency. About the Role We are looking for a technically deep TPM to independently define and lead multiple programs across data platforms, online databases and storage infrastructure. You will connect model, product and data-consumer requirements to architecture, and work with the relevant engineering teams to take new capabilities through production adoption and repeatable expansion. The design scope is exabyte-scale storage and infrastructure spanning multiple millions of CPU cores. The challenge is not simply forecasting more resources: it is making complete, workload-ready capacity repeatable, with a clear path from product requirements through architecture, deployment and validation. A data pipeline, database query, file operation or execution snapshot can affect whether a product or agent succeeds; you will connect those outcomes to the systems underneath. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Translate model, product and data-platform needs into precise access patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency. Partner with engineering to transform data and storage architecture into repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together.

AWSAzureRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan

SQLAWSRestAI
MT
📍 Richardson, TX, United States
✓ Quality checkedCompany trend +1216.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. We are seeking a Senior Analog Design Engineer to join our High Bandwidth Memory (HBM) team. This role is responsible for the architecture, design, simulation, and silicon bring ‑ up of high ‑ performance analog and mixed ‑ signal circuits used in HBM PHYs and supporting infrastructure. The ideal candidate has deep expertise in high ‑ speed I/O, clocking, power management, and advanced ‑ node analog design, and has successfully delivered silicon to production. Responsibilities will include, but are not limited to: Design and own critical HBM analog circuits, including: High ‑ speed transmitters and receivers Clock generation and distribution (PLLs, DLLs, CDRs) SerDes ‑ related analog blocks Biasing, reference, and calibration circuits </

AIRecruitment
🔔

Get new infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime