Jobs in United States

Infrastructure Engineer in United States

1,475 active opportunities · Updated October 2026

Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.8%

From $128K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: The Strategic Sourcing team is part of Airbnb’s Global Procurement organization and plays a key role in enabling Airbnb to operate efficiently and scale sustainably. We partner closely with internal stakeholders to design and deliver sourcing strategies that align with Airbnb’s business priorities while driving innovation, value, and stakeholder success. We are a collaborative, forward-thinking team that embraces an AI-first mindset — leveraging technology, automation, and data to simplify processes and unlock new efficiencies. Guided by Airbnb’s mission to create a world where anyone can belong, we focus on delivering value through curiosity, creativity, and integrity. Key cross-functional partners: Finance, Legal, Business Operations. The Difference You Will Make: As the Sourcing Manager supporting Engineering and Infrastructure, you will partner directly with engineering leaders to shape how Airbnb sources and manages critical technology capabilities. This role sits at the center of major technical investments, from infrastructure platforms to emerging AI technologies. You will design sourcing strategies that balance cost efficiency, engineering velocity, and long-term scalability. You will also play a key role in advancing how sourcing integrates AI and automation into its workflows. Rather than simply executing procurement processes, you will help modernize them by improving how market intelligence, supplier comparisons, and negotiation insights are generated and used in decision-making. Success in this role means helping engineering teams make faster, better technology d

T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing

$100K – $500K/yr

Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent builds AI and RISC-V compute hardware that requires best-in-class design tools and intellectual property to compete globally. The Sr. Strategic Sourcing Manager, Engineering Infrastructure leads procurement strategy for the design ecosystem, owning vendor negotiations, IP licensing agreements, and tool sourcing that enable our worldwide engineering teams to innovate at speed. The role suits someone who takes ownership of complex, multidimensional deals, operates independently with minimal oversight, and commands credibility with vendor executives, internal stakeholders, and the C-suite. This role is hybrid or remote, based out of Santa Clara, CA; Austin, TX or Toronto, ON. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A strategic negotiator with 5+ years of experience sourcing complex technology (IP, design tools, or semiconductor components) in fast-paced environments Someone who builds trust-based vendor relationships and leverages them to unlock better terms, faster delivery, and innovative solutions A detail-oriented operator who thrives managing multiple concurrent deals while maintaining executive visibility and cross-func

AWSAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -86%

$196K – $261K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role: Millions of people use Notion — and this number is increasing every day. That means millions of people trust us to deliver a fast, reliable, and secure experience, and we value this more than anything. We want to keep earning trust, while also continuing to amaze our users with the tools they can build in Notion. The Product Analytics Platform team owns the foundational systems that turn Notion's raw product signals into reliable and actionable insights at scale: how we instrument events, experiment and roll out features, define and trust metrics, and build analytics surfaces that interact with AI agents.. This is an early, high-leverage moment. You’ll join a newly formed team of talented engineers in our mission to build the best in class foundational platforms that are key to the company’s business and product. What You'll Achieve: You'll help shape a brand-new platform team and have real influence over the architecture and roadmap. You'll own critical infrastructure that spans experimentation, feature gating, event logging, schematization, governance and play a pivotal role in the development of systems, tools and

TypeScriptNode.jsCI/CDRest
D
📍 New York, New York, United States
✓ High-confidence listingCompany trend -84.7%
Quick readStrong listing-quality and freshness signals

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at a high scale - trillions of data points per day — providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity We are looking for an experienced software engineer to join our CI/CD Security Team within our SDLC Security organization. We work at the intersection of security and engineering infrastructure to secure Datadog's continuous integration and continuous delivery systems. Our responsibilities include hardening pipelines, protecting credentials, and enforcing tightly scoped access controls. We also develop authorization and verification mechanisms to ensure that only trusted code and approved processes can reach production. In this role, you will shape and build a new security layer for our CI/CD infrastructure and drive its adoption across the engineering organization. You will solve challenging systems problems around trusted build provenance, secure secret delivery, and real-time policy enforcement at high throughput. The work sits directly in the critical path of software delivery, where strong security guarantees have to coexist with low latency, high reliability, and a seamless developer experience. You’ll join at an ideal time to make a big impact, as the need for robust software supply chain security is higher than ever. Datadog is growing rapidly, and AI-assisted development is increasing both the pace of software delivery and the amount of activity flowing through our CI/CD systems. Securing that scale without slowing engineers down requires strong software engineering fundamentals, thoughtful automation, and security controls designed to operate reliably at high throughput. At Datadog, we pla

JavaScriptPythonJavaAWS
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $399.4K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The platforms in our Engineering Acceleration org dictate how thousands of Roblox engineers ship, and how every backend service gets built, deployed, and kept reliable at scale. They sit on the critical path for production reliability, developer velocity, and security posture. Reporting to Andrew Swerdlow, you will rethink the entire engineering toolchain from the ground up to be agentic-native, creating a world where AI agents are first-class participants in software development and humans set direction and supervise. This is a unique opportunity to define what modern engineering infrastructure looks like at one of the largest platforms in the world. You will: Reimagine the engineering toolchain as agentic-native, designing the platforms, guardrails, and feedback loops that let AI agents safely drive migrations, validation, and routine operational work. Set a bold technical direction for AI-driven quality, including agent-generated tests, automated coverage of untested paths, intelligent verification, and continuous-deployment workflows. Make software quality and SEV prevention a measurable property of the platform by investing in safe-change mechanisms, automated verification, progressive

AWSKubernetesCI/CDGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer you will join our OpenAI engineers and researchers in building, operating and securing transformational AI technologies. This role will focus on all aspects of Detection & Response but with a strong emphasis on detecting insider threats and influencing controls to safeguard OpenAI's most sensitive assets. In this role, you will: In this role, you will: Innovate on Detection and Response infrastructure to engineer and automate end-to-end detection and investigation workflows. Develop, measure, and tune detection rules to ensure effective and sustainable operations. Drive projects across OpenAI’s technology stack with a focus on insider threats, ranging from access abuse and intellectual property theft to novel risks emerging within AI infrastructure. Partner closely with cross-functional stakeholders, including HR, Legal, and peer investigative teams, providing technical expertise and evidence to support investigations. Collaborate on cutting-edge AI research, and use AI to improve OpenAI’s Security posture. You might thrive in this role if you: 5+ years experience working in a detection/response or insider-risk role.. We are seeking mid-level and senior candidates. You have broad familiarity with operating systems and platforms such as macOS, Windows, Linux, and Kubernetes, along with experience in cloud infrastructure. Knowledge of modern adversary tactics and attack paths, data exfiltration techniques, and h

PythonAWSKubernetesLinux
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Artifacts team is building the AI-native creation layer for documents, spreadsheets, slide decks, dashboards, reports, analyses, and new forms of interactive work products. We are rethinking what creation looks like when models can move from an ambiguous user goal to a polished, editable artifact with strong structure, taste, correctness, and speed. This is a high-agency team working across product, infrastructure, and research. We partner closely with model training teams to shape how frontier models create artifacts, and with ChatGPT product teams to turn those capabilities into experiences that millions of people can use. The work spans full-stack product engineering, model integration, rendering and editing systems, collaboration, storage, evaluation loops, and production reliability. Our ambition is to build the premier product experience for AI-generated artifacts: starting with familiar work products like slides, sheets, and docs, then expanding into new artifact types that are only possible in an AI-native world. About the Role As Engineering Manager, Artifacts, you will lead and grow the engineering team responsible for building this product and technical foundation. You will manage a team of full-stack and infrastructure-oriented engineers, set technical direction, and stay hands-on enough to shape architecture and debug hard problems. This role sits at the intersection of product engineering, research, and infrastructure. You will partner with researchers on how models are trained and evaluated for artifact creation, with product and design on the user experience. This is a strong fit for a technical manager who wants to build and ship, not only coordinate. The team has a fast trajectory, so you will help define both the product surface and the team that builds it. In this role, you will: Lead, manage, and grow a team building AI-native artifact creation experiences across documents, spreadsheets, slide decks, and emerging artifact form

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our San Francisco office. We are a hybrid environment that combines the energy and connect

PythonJavaSQLAWS
B
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our New York office. We are a hybrid environment that combines the energy and connections

PythonJavaSQLAWS
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $154K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

PythonKubernetesAISwift
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Technical Support team is responsible for ensuring that developers and enterprises can reliably build mission critical solutions using OpenAI models. We provide technical guidance, resolve complex issues and support customers in maximizing value and adoption from deploying our highly-capable models. We work closely with Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. We think from an automation-first mindset and leverage the latest in AI to scale our support operations. Join the Senior Support Engineering (SSE) team at OpenAI and help shape the future of Technical Support in the age of AI. About the Role We are looking for a Senior Support Engineer to collaborate directly with our strategic enterprise accounts and product teams, helping solve some of the most difficult problems faced by our Customers. You will be part of the best technical troubleshooting team at OpenAI, and our Customers and Engineering teams will look to you for technical guidance in addressing the most technically difficult issues in our environment. As a Senior Support Engineer, you will design and run operational processes to monitor our top strategic customers and a 24x7 response team. You’ll work closely with our Infrastructure and Engineering teams to deliver the best possible experience to customers at scale. Working directly with our most strategic Customers - You will be crucial to the success of the most innovative, disruptive, and high-scale AI solutions being built with the OpenAI API platform. The nature of this role will be low volume, high difficulty. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Be among the foremost technical and troubleshooting experts for our API platform at OpenAI. You are the last line of defense before the core Engineering team. Proactively iden

PythonAWSRestAI
T
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. This role sits at the center of cutting-edge AI hardware development, keeping the servers, PCIe systems, and engineering infrastructure running that power next-generation compute. You’ll be hands-on with rapidly evolving prototype and production systems, installing, maintaining, and troubleshooting hardware in fast-paced R&D and data center environments. Acting as a critical bridge between hardware engineers, software teams, and IT, you’ll help ensure seamless access to the platforms that turn ideas into working silicon and systems. This role is onsite, based out of Toronto, Canada or Austin, Texas. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are A hands-on hardware professional who enjoys building, maintaining, and troubleshooting complex computing systems. Comfortable working in fast-paced R&D environments where hardware, firmware, and software are constantly evolving. Knowledgeable in computer architecture, operating systems, and hardware diagnostics, with strong problem-solving skills. Collaborative, detail-oriented, and motivated to improve processes through documentation, scripting, and automation. What We Need Inst

AWSAISEMHR
P
📍 New York, NY, United States
✓ High-confidence listing

$175K – $275K/yr

Quick readStrong listing-quality and freshness signals

A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Develop and maintain the information security policy and standards library, aligning it with business priorities, regulatory expectations, and control objectives Lead independent assessments of the information security program, including regulatory examinations and third-party evaluations Identify technology risks across a complex business environment and drive mitigation aligned with our firm’s standards and control expectations Investigate data privacy inquiries and privacy-related events, assess business impact, and coordinate timely response activities Partner with technology, legal, compliance, and business stakeholders to translate risk findings into practical remediation plans Advise control owners on policy interpretation, risk treatment, and evidence expectations for assessments and examinations Prepare clear reports for management on team activity, emerging risks, remediation progress, and key decisions Maintain accurate risk, policy, priv

Artificial IntelligenceAI
P
📍 New York, NY, United States
✓ High-confidence listing

$170K – $250K/yr

Quick readStrong listing-quality and freshness signals

A Career with Point72’s Technology Team As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications. As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business. What you’ll do Optimize cloud financial operations to maximize value from cloud investments, including rapidly growing artificial intelligence (AI) and machine learning workloads Provide actionable insights on cloud spend, SaaS license optimization, and emerging AI cost drivers, including model inference and usage-based consumption Implement tooling, tagging standards, and processes that improve cost visibility and optimization across cloud, SaaS, and AI workloads Monitor large language model API consumption and GPU-intensive infrastructure to identify cost trends, anomalies, and optimization opportunities Build financial models to forecast cloud, SaaS, and AI expenditures for budgeting cycles, commitment decisions, and vendor negotiations Design cost allocation, tagging, showback, and chargeback models that attribute spend to the teams, applications, and use cases driving it Educate engineering and business owners on cloud financial management practices th

AWSAzureMachine LearningArtificial Intelligence
🔔

Get new infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime