Jobs in United States

Senior Infrastructure Engineer in United States

2,140 active opportunities · Updated October 2026

Explore current senior infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

I
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -92.5%

From $265K/yr

Quick readStrong listing-quality and freshness signals

We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Instacarts Data Infrastructure organization builds and operates the systems that power our company’s data ecosystem, including a modern open data lakehouse on Apache Iceberg, a multi-engine compute platform for stream and analytical workloads, and self-serve tooling that helps Product, Data Science, ML, Ads, Finance, and engineering teams move fast with data. We’re looking for a Staff Software Engineer, Data Infrastructure to join our Data Governance and Foundations Team. In this role, you’ll serve as a senior technical leader owning the architecture and delivery of our open lakehouse foundation, governance and access patterns, and multi-engine compute strategy—balancing today’s reliability with the next three to five years of scale, maturity, and cost efficiency. You’ll collaborate closely with engineering leadership and stakeholders across Data Science,

PythonSQLAWSAI
A
📍 United States· Full-time
✓ High-confidence listingCompany trend -98.9%

From $232K/yr

Quick readStrong listing-quality and freshness signals

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: Airbnb's mission is to create a world where people can Belong Anywhere. As we grow to achieve that mission, we're looking to add a technical, hands-on, and mission-driven Senior Manager of Technical Program Management (TPM) for the Infrastructure team to lead AI Platforms initiatives spanning Search Relevance, Machine Learning Infrastructure (MLI), and Search Knowledge Infrastructure (SKI). This person will partner with our engineering & product teams to build/enhance data-driven decision making across multiple areas in Airbnb. They will partner closely with senior ICs and engineering directors across Airbnb on processes, strategy and execution. In this role, you will collaborate closely with the Machine Learning (ML) engineers, infrastructure engineers and product managers from across Airbnb to develop holistic solutions that ensure a vibrant and equitable marketplace. The Search Relevance, Machine Learning Infrastructure (MLI), and Search Knowledge Infrastructure (SKI) teams power the ranking, recommendations, model infrastructure, and search intelligence capabilities that make up a core part of Airbnb's AI Platforms, including: Search Ranking for Stays, Experiences, and Services, Home Page Recommendations, Search Bar AutoSuggest and Autocomplete, Marketing Technology, and the ML infrastructure and platform capabilities that support them. The team also creates the fabric for personalization and the underlying model infrastructure, used both on and off platform for Airbnb. We collaborate with many teams including Search Infrastructure, Guest and Host Experiences, MD

Machine LearningAIGoRust
C
📍 Work From Home, United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is seeking a highly skilled Staff Data Engineer, Observability Engineering to join the Enterprise Observability Platform organization and help advance the next generation of observability, infrastructure, and security data capabilities. The Staff Data Engineer, Observability Engineering will play a critical role in designing, building, and operating scalable data pipelines and data products that power enterprise observability, operational intelligence, and security analytics across the organization. The Staff Data Engineer, Observability Engineering is a senior individual contributor responsible for developing and optimizing Databricks-based data engineering solutions that ingest, transform, govern, and deliver high-volume telemetry, infrastructure, application, and security data. This role combines deep hands-on technical execution with ownership of engineering excellence, operational reliability, performance optimization, and data platform best practices. Working closely with Observability Engineering, Security Engineering, Infrastructure Engineering, and Data Platform teams, the Staff Data Engineer, Observability Engineering will contribute to the evolution of the enterprise observability lakehouse by building resilient ingestion frameworks, establishing data quality standards, enhancing governance controls, and driving efficient, scalable data processing patterns. The id

PythonSQLAzure
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

AWSRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $233.6K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Security Software Engineer for Infrastructure Security you will be a part of the Information Security organization and report to the Senior Manager of Infrastructure Security. You will help shape the future of Platform Security at Roblox. We work closely with Production IAM, Network Security, and Cloud Security at Roblox. You Will: Identify security gaps and threats in our cloud and on premise infrastructure, partnering with Governance Risk and Compliance teams to create standards and policies along the way. This will help Roblox meet regulatory and compliance requirements. Harden our infrastructure by introducing secure by default configurations, designs and guardrails for all developers at Roblox. Own and drive solutions that enable Roblox engineers to design, build, and use infrastructure securely at scale. Work closely with other InfoSec teams (AppSec, D&R, GRC, CorpSec, CloudSec, NetSec) and partner with engineering teams across Roblox, specifically the Infrastructure organization, to ensure the secure outcomes of security and product driven initiatives. You Have: 5+ years of experience writing code and/or relevant technical experience. Experience with

AWSKubernetesGitLinux
G
📍 United Kingdom; Remote, United States· Full-time· Remote
✓ High-confidence listingCompany trend -100%

From $126.4K/yr

Quick readStrong listing-quality and freshness signals

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure. This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience

AWSGCPKubernetesCI/CD
D
📍 New York, New York, United States
✓ High-confidence listingCompany trend -88.2%
Quick readStrong listing-quality and freshness signals

About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at a high scale - trillions of data points per day — providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity We are looking for an experienced software engineer to join our CI/CD Security Team within our SDLC Security organization. We work at the intersection of security and engineering infrastructure to secure Datadog's continuous integration and continuous delivery systems. Our responsibilities include hardening pipelines, protecting credentials, and enforcing tightly scoped access controls. We also develop authorization and verification mechanisms to ensure that only trusted code and approved processes can reach production. In this role, you will shape and build a new security layer for our CI/CD infrastructure and drive its adoption across the engineering organization. You will solve challenging systems problems around trusted build provenance, secure secret delivery, and real-time policy enforcement at high throughput. The work sits directly in the critical path of software delivery, where strong security guarantees have to coexist with low latency, high reliability, and a seamless developer experience. You’ll join at an ideal time to make a big impact, as the need for robust software supply chain security is higher than ever. Datadog is growing rapidly, and AI-assisted development is increasing both the pace of software delivery and the amount of activity flowing through our CI/CD systems. Securing that scale without slowing engineers down requires strong software engineering fundamentals, thoughtful automation, and security controls designed to operate reliably at high throughput. At Datadog, we pla

JavaScriptPythonJavaAWS
B
📍 San Francisco, California, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our San Francisco office. We are a hybrid environment that combines the energy and connect

PythonJavaSQLAWS
B
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$192K – $240K/yr

Quick readStrong listing-quality and freshness signals

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our New York office. We are a hybrid environment that combines the energy and connections

PythonJavaSQLAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team The Technical Support team is responsible for ensuring that developers and enterprises can reliably build mission critical solutions using OpenAI models. We provide technical guidance, resolve complex issues and support customers in maximizing value and adoption from deploying our highly-capable models. We work closely with Technical Success, Product, Engineering and others to deliver the best possible experience to our customers at scale. We think from an automation-first mindset and leverage the latest in AI to scale our support operations. Join the Senior Support Engineering (SSE) team at OpenAI and help shape the future of Technical Support in the age of AI. About the Role We are looking for a Senior Support Engineer to collaborate directly with our strategic enterprise accounts and product teams, helping solve some of the most difficult problems faced by our Customers. You will be part of the best technical troubleshooting team at OpenAI, and our Customers and Engineering teams will look to you for technical guidance in addressing the most technically difficult issues in our environment. As a Senior Support Engineer, you will design and run operational processes to monitor our top strategic customers and a 24x7 response team. You’ll work closely with our Infrastructure and Engineering teams to deliver the best possible experience to customers at scale. Working directly with our most strategic Customers - You will be crucial to the success of the most innovative, disruptive, and high-scale AI solutions being built with the OpenAI API platform. The nature of this role will be low volume, high difficulty. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Be among the foremost technical and troubleshooting experts for our API platform at OpenAI. You are the last line of defense before the core Engineering team. Proactively iden

PythonAWSRestAI
G
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $154K/yr

Quick readStrong listing-quality and freshness signals

Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.​ This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s

PythonKubernetesAISwift
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $221.4K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. WHY DATA SCIENCE & ANALYTICS? The Data Science & Analytics organization’s mission is to increase our speed, frequency and acumen of making decisions at scale by instilling a data-influenced approach to building products. We cover a wide area of the data spectrum including analytical data engineering, product analytics, experimentation, causal inference, statistical modeling and machine learning. Aligned and partnering with product verticals, we use this extensive toolbelt to discover new opportunities and unmet use cases, influence and shape the product roadmap and prioritization, build data products and measure impact on our community of players and creators. WHY CREATOR SERVICES? At Roblox, the Creator Services team enables unbounded creation through reliable core services and novel AI applications. As a Senior Data Scientist focused on Machine Intelligence , you will bridge the gap between high-tier engineering infrastructure and cutting-edge ML applications. You will be the primary strategic partner to product and engineering leadership, transforming unstructured data into actionable business insights and user-facing products. This is a "zero-to-one" environment. You will be tas

AWSGitMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role As a Security Engineer you will join our OpenAI engineers and researchers in building, operating and securing transformational AI technologies. This role will focus on all aspects of Detection & Response but with a strong emphasis on detecting insider threats and influencing controls to safeguard OpenAI's most sensitive assets. In this role, you will: In this role, you will: Innovate on Detection and Response infrastructure to engineer and automate end-to-end detection and investigation workflows. Develop, measure, and tune detection rules to ensure effective and sustainable operations. Drive projects across OpenAI’s technology stack with a focus on insider threats, ranging from access abuse and intellectual property theft to novel risks emerging within AI infrastructure. Partner closely with cross-functional stakeholders, including HR, Legal, and peer investigative teams, providing technical expertise and evidence to support investigations. Collaborate on cutting-edge AI research, and use AI to improve OpenAI’s Security posture. You might thrive in this role if you: 5+ years experience working in a detection/response or insider-risk role.. We are seeking mid-level and senior candidates. You have broad familiarity with operating systems and platforms such as macOS, Windows, Linux, and Kubernetes, along with experience in cloud infrastructure. Knowledge of modern adversary tactics and attack paths, data exfiltration techniques, and h

PythonAWSKubernetesLinux
O
📍 San Francisco, California, United States
✓ High-confidence listingCompany trend 0%
Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Okta is the identity standard. The Okta Platform is an independent and neutral platform that securely connects the right people to the right technologies at the right time. We help organizations do two things - secure and manage their extended enterprise, and transform their users’ experiences. Okta's Core Engineering team is responsible for building and evolving shared infrastructure and services that lay the foundation for what other engineering teams build on. We're in charge of common shared services like search, cache, configuration management, frameworks for async job management, and email pipeline, to name a few. We're cloud native, where redundancy, multi-tenancy, scale, resource optimization and resiliency are first class citizens. With Okta's mantra of 'Always On!' there's never a dull moment. Our biggest asset is our team of passionate engineers and technically minded managers. We're looking for a staff level backend engineer to join a team of highly skilled and talented team players who're proud of what they own and deliver. Our elite team is fast, creative and flexible; with a weekly release cycle and individual ownership we expect great things from our engineers and reward them with stimulating new projects, new technologies and the chance to have significant equity in a company that is changing the cloud computing landscape forever. You will: Work with engineering teams to design, develop and deliver cloud based infrastructu

JavaRedisAWSDocker
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a member of the Infrastructure Foundation Hardware Engineering team, you will play a key role in enabling our mission to deliver a reliable, high-performing, and cost-efficient infrastructure that powers the world’s play. In this specialized role, you will be the technical lead for our GPU and AI accelerator ecosystem. You will be responsible for the full lifecycle of GPU hardware, from initial architectural evaluation and firmware qualification to large-scale fleet integration and performance tuning. You will ensure that Roblox’s massive-scale rendering and ML workloads run on the most optimized and stable hardware possible. You Will: Architect & Prototype: Prototype next-generation GPU-accelerated hardware platforms, ensuring seamless integration between high-density compute nodes, high-speed interconnects (NVLink/PCIe Gen5/6), and system firmware. GPU Optimization: Drive the integration, performance testing, and debugging of GPUs in our fleet, focusing specifically on hardware-level optimizations, driver tuning, and thermal/power management. Validation & Certification: Develop and execute rigorous evaluation and stress-testing strategies for GPU-heavy server platforms to ensur

PythonAWSGitLinux
🔔

Get new senior infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime