Jobs in United States

Senior Infrastructure Engineer in United States

2,189 active opportunities · Updated October 2026

Explore current senior infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $326.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when

PythonJavaAWSGit
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Join Roblox as an Engineering Manager of Application Security and lead a team responsible for improving the security of our products, services, and development ecosystem. In this role, you will drive security across the software lifecycle, partnering with engineering teams to identify risks, improve secure development practices, and build scalable solutions that protect Roblox at scale. You will balance hands-on security work with longer-term investments in automation, tooling, and developer enablement. You will work closely with engineering, infrastructure, and security teams to reduce risk while enabling teams to move quickly and safely. This role reports to the Senior Manager of Application Security and is based in San Mateo with a hybrid schedule. You Have: 8+ years of experience in Information Security 2+ years of experience managing engineers Strong background in Application Security or Product Security Experience driving security programs across the software development lifecycle Solid understanding of common vulnerabilities (e.g., OWASP Top 10) and secure coding practices Experience working closely with engineering teams in modern environments (cloud, microservices, CI/CD) Prov

AWSCI/CDGitMicroservices
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team The Solutions Engineering team is made up of trusted technical advisors who help organizations adopt OpenAI’s technology safely, effectively, and responsibly. We partner closely with customers, Sales, Product, Engineering, and Security to translate frontier AI capabilities into practical workflows that create real-world impact. Cybersecurity is one of the most urgent areas where AI can help. As frontier models become more capable at reasoning over code, logs, infrastructure, and security evidence, organizations need guidance on how to evaluate, validate, and deploy these systems safely. Our goal is to help customers move from identifying risks to implementing solutions. About the Role We are committed to bringing together people from diverse backgrounds and perspectives who are excited to help build and deploy safe, useful AI. We are seeking a solutions engineer to partner with our Enterprise customers and ensure they achieve tangible business value from our models through the OpenAI suite of products. You will partner with senior business stakeholders to understand their pre-sales needs, guide their AI strategy, and identify the highest value use cases and applications. You will work with business and technical teams to demonstrate the value of our solutions and recommend architectural patterns to kickstart their implementation and development. You will work closely with Enterprise Sales, Security, and Product teams. We are looking for a Field Security Specialist to help security leaders and hands-on practitioners understand how OpenAI models, APIs, Codex, and agentic workflows can be applied to real cybersecurity use cases. This is a customer-facing specialist role for someone who can move fluidly between CISO-level conversations, practitioner-level technical depth, and hands-on solution design. You’ll help customers evaluate OpenAI for workflows like secure code review, vulnerability triage, threat modeling, remediation, SOC workflows, detection en

AWSCI/CDGitRest
C
📍 Tampa Florida United States, United States
✓ Quality checkedCompany trend +800%

At Citi Services - Global Trade and Working Capital Solutions (TWCS) Technology Organization, we are on a mission to harness the power of data to drive innovation, create exceptional customer experiences, and solve complex business challenges. Our data team is at the heart of this mission, building the scalable and resilient infrastructure that turns data into our most asset. We are a passionate, collaborative group dedicated to pushing the boundaries of what's possible. The Opportunity We are seeking an Engineering Director to join our development team. The ideal candidate is a seasoned technologist with extensive experience in building and delivering data & AI solutions for business functions. This individual will be directly and fully accountable for solution delivery and should have a history of creating strategic technology architecture roadmaps aligned with business outcomes. The successful candidate will define and execute the technology roadmap for the Data & AI portfolio of TWCS, providing strategic direction and critical input into technology decisions to ensure a scalable buildout. This role involves building strong relationships with senior business and technology partners, driving agile execution, and leading a team of expert engineers to deliver with velocity and quality. Applicants should demonstrate exceptional technical acumen, a strong data engineering background, and a proven ability to lead and provide technical direction to high-performing engineering teams. A proven expertise in Data Engineering, Data Analytics, and AI Engineering delivery at scale is essential. Responsibilities: Manage/develop multiple teams of professionals to accomplish established goals and conduct personnel duties for team (e.g. performance evaluations, hiring and disciplinary actions) as well as ensure team adheres to best practices and processes Develop vision for team aroun

PythonJavaAWSAzure
MT
📍 Richardson, TX, United States
✓ Quality checkedCompany trend +1150%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. We are seeking a Senior Analog Design Engineer to join our High Bandwidth Memory (HBM) team. This role is responsible for the architecture, design, simulation, and silicon bring ‑ up of high ‑ performance analog and mixed ‑ signal circuits used in HBM PHYs and supporting infrastructure. The ideal candidate has deep expertise in high ‑ speed I/O, clocking, power management, and advanced ‑ node analog design, and has successfully delivered silicon to production. Responsibilities will include, but are not limited to: Design and own critical HBM analog circuits, including: High ‑ speed transmitters and receivers Clock generation and distribution (PLLs, DLLs, CDRs) SerDes ‑ related analog blocks Biasing, reference, and calibration circuits </

AIRecruitment
B
📍 Berkeley, Macau S.a.r., United States
✓ Quality checkedCompany trend -20.1%

Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <

AWSAzureDockerKubernetes
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for an experienced Executive Assistant to support our CTO (co-founder) and Head of Engineering. This is a highly operational role that goes well beyond calendar management. You’ll own the day-to-day operating rhythm of the Engineering organization, ensuring leaders are prepared, priorities stay coordinated, and critical meetings, communications, and follow-ups happen seamlessly. You’ll partner closely with senior engineering leaders and serve as a trusted point of coordination for employees, customers, candidates, and external partners. Success in this role comes from exceptional organization, judgment, attention to detail, and the ability to keep many moving pieces aligned in a fast-growing environment. RESPONSIBILITIES Own complex calendar management for the CTO and Head of Engineering, balancing shifting priorities while ensuring time is allocated intentionally Ensure leaders are prepared for every day and every meeting by proactively managing agendas, materials, context, logistics, and follow-ups so time is used effectively and decisions move forward Support forward-looking calendar planning, coordinating recurring operating cadences including roadmap planning, leadership meetings, P0 reviews, Engineering All Hands, and other cross-functional forums Own the operational cadence of the Engineering organization, including weekly leadership meetings, monthly Show & Tells, Engineering All Hand

Machine LearningAIGoRust
C
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.2%

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role Is Different This is not a typical “Applied Scientist” or “ML Engineer” role. As a Member of Technical Staff, Applied ML, you will: Work directly with enterprise customers on problems that push LLMs to their limits. You’ll rapidly understand customer domains, design custom LLM solutions, and deliver production-ready models that solve high-value, real-world problems. Train and customize frontier models — not just use APIs. You’ll leverage Cohere’s full stack: CPT, post-training, retrieval + agent integrations, model evaluations, and SOTA modeling techniques. Influence the capabilities of Cohere’s foundation models. Techniques, datasets, evaluations, and insights you develop for customers will directly shape the next generation of Cohere’s frontier models. Operate with an early-startup level of ownership inside a frontier-model company. This role combines the breadth of an early-stage CTO with the infrastructure and scale of a deep-learning lab. Wear multiple hats, set a high technical bar, and define what Applied ML at Cohere becomes. Few roles in the industry combine application, research, customer-facing engineeri

PythonGitAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C

SQLAWSAzureGCP
G
📍 California, El Segundo, United States
✓ High-confidence listingCompany trend -1.6%

$115.3K – $149.2K/yr

Quick readStrong listing-quality and freshness signals

We’re here for one reason and one reason only – to cure cancer. Every moment is dedicated to developing treatments and every action moves us one step closer to our goal. We’ve made incredible scientific breakthroughs and our pioneering personalized CAR T-cell therapies have changed the paradigm. But we're not finished yet. Join Kite, as we make even bigger advances in cancer therapies, and help shape where our business and medical science goes next. We believe every employee deserves a great leader. People Leaders are the cornerstone to the employee experience at Gilead and Kite. As a people leader now or in the future, you are the key driver in evolving our culture and creating an environment where every employee feels included, developed and empowered to fulfil their aspirations. Join Kite and help create more tomorrows. Job Description Position Summary The Senior IT Quality Engineering Specialist serves as the technical lead for the Kite Laboratory Information Management System (KLIMS) within North America West Coast (El Segundo). This role is responsible for the operational stability, compliance, maintenance, enhancement, and technical governance of the LabVantage platform and associated integrations. The position partners closely with Quality Control, Manufacturing, Validation, Infrastructure, and Kite Business stakeholders to ensure reliable and compliant laboratory operations. Key Responsibilities System Administration & Operational Support • Provide day-to-day administration and technical support for KLIMS and related

JavaSQLRecruitment
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $128K/yr

Quick readStrong listing-quality and freshness signals

We’re looking for an experienced sourcing leader to join the Datadog Procurement team and help grow the Strategic Sourcing group. Make an impact by partnering closely with senior technology and business leaders, owning our fastest-growing AI, neocloud, and inference spend end-to-end, and continuing to prove the value that Strategic Sourcing brings to the organization. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the AI spend category end-to-end, covering foundation-model and AI APIs, AI development and productivity tools, and AI-enabled SaaS (with neocloud and inference compute as an emerging area), while developing and executing category strategies aligned to business objectives Champion AI and automation adoption across the sourcing function by identifying tools (e.g. Claude, ChatGPT) and building repeatable workflows that make sourcing faster, smarter, and more scalable Lead high-value AI vendor negotiations spanning foundation-model and API agreements, AI development and productivity tools, and AI SaaS, structuring seat- and consumption-based pricing across net-new purchases and strategic renewals, and building neocloud and GPU-capacity capability as the category grows Partner with Engineering and Finance to turn architectural and consumption trade-offs into commercial business cases, forecasts, and savings targets Build pricing and consumption models and apply FinOps discipline to quantify buying scenarios and uncover savings across a fast-moving spend base Work alongside executive and senior technology leadership as a trusted advisor on AI and neocloud investment decisions, influencing strategy and commercial trade-offs at the leadership level Track category KPIs (savings, pipeline, cycle time), monitor AI market, vendor, and pricing trends,

R
📍 Goodyear, Arizona, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $185K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Manager, Data Center Operations, you'll help us scale our Core Data Center and hardware infrastructure at a time of incredible growth for our business. At Roblox, you'll have boundless opportunities to shape the future of the Imagination Platform™ and demonstrate your passion for delivering thoughtful solutions in front of a global audience. If you know what it takes to build and operate hardware infrastructure that can sustain millions of concurrent players year-round and you take play as seriously as we do, you'll fit right into our highly experienced and ever-expanding engineering team. You will report to the Senior Manager of Data Center Operations. This will be a position based in Goodyear, AZ. You will: Develop and maintain the Core Data Center and hardware infrastructure to meet the large-scale and real-time requirements of our Imagination Platform™ to ensure our community has an awesome experience anywhere in the world. This includes all aspects of the server, network infrastructure, power, and environmental monitoring. Lead a growing team of data center engineers focusing on rack deployments, hardware troubleshooting and break-fix, and decommissioning. Identify and solve criti

AWSGitAIGo
S
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.6%

$200K – $240K/yr

Quick readStrong listing-quality and freshness signals

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role The Platform org at Sentry is the engine that everything else runs on, spanning developer platform, core infrastructure underpinning the product, SRE, security, and IT. It's a broad, technically complex, and deeply consequential organization. We're looking for a Staff Technical Program Manager who can operate at the intersection of technical depth and strategic execution: someone who thrives in complexity, builds trust with senior engineering leaders, and has a gift for turning ambiguity into clarity and momentum. You'll report to the Head of Technical Program Management and work closely with the VP of Platform Engineering, their staff, and partner teams across Engineering, Product, and Design (EPD). This is a high-visibility role with real influence. You'll work directly with the CTO and senior leaders, shape how the Platform org operates, and help Sentry scale through one of its most important chapters. In this role, you will: Drive strategic execution. Partner with the VP of Platform Engineering and senior EPD leaders to translate priorities into clear, measurable plans. Own sequencing, milestones, and key decisions, and make sure leadership always has reliable visibility into progress and tradeoffs. Build data-driven delivery health. Establish the metrics and dashboards that reflect delivery confidence, risk, and engineering health across the Platform org. Keep planning and reporting high-signal and lightweight, focused on outcomes rather than activity. Manage capacity, dependencies, and risk. Create visibility into resourcing, cross-team dependencies, and constraints so leaders can align investment to the

E(
📍 San Francisco Bay Area, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

About Ema Ema is building the world’s leading Agentic AI platform to transform enterprise productivity. We enable organizations to delegate repetitive tasks to Ema, the Universal AI Employee, delivering 10x gains in workforce efficiency, across functions. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production — we ship real systems that run real business processes at scale. The residency You own one hard problem end to end. You write the proposal, build the system, design the evaluation, ship behind a gate, and finish with a write-up of what turned out to be true, including the parts that didn't work. You'll sit in the production codebase with a senior mentor and real production data. Recent residents have shipped self-improving harnesses, inference-cost work, agent memory, and eval infrastructure. Your project gets scoped with you, not handed to you. The problem space The loop we care about: production traces become data, data becomes training and evaluation, and better agents produce better traces. Projects live somewhere on that loop. Harness and inference-time work. Context engineering, tool and skill design, orchestration, and deciding where extra inference compute actually pays. Self-improvement loops run behind hard fences. Post-training for agents. SFT on curated trajectories, preference optimization, RL on real agent tasks. Reward design where outcomes are verifiable, process vs. outcome supervision, distilling frontier behavior into cheaper models. Environments and rewards. Turning enterprise workflows into training and eval environments: fi

PythonAIGo
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

AIGoExcelSEM
🔔

Get new senior infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime