Jobs in United States

Senior Infrastructure Engineer in United States

2,189 active opportunities · Updated October 2026

Explore current senior infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product workloads. The company seeks a Senior Technical Program Manager (TPM) to lead complex, cross-functional programs powering NVIDIA’s next-generation AI software platforms. In this role, you will drive software initiatives across platform services, cloud infrastructure, and system integration. The focus is on enabling scalable, reliable, and supportable software for AI workloads. You will be responsible for managing high-impact engineering programs within a dynamic, fast-paced roadmap, aligning priorities across teams, and ensuring timely, high-quality delivery. This role requires strong technical competence, a proactive approach, and the ability to operate effectively across multiple levels of the organization. This is a software-first TPM role. The ideal candidate has extensive experience managing software initiatives. They also understand the full-stack environment, including infrastructure dependencies, system bring-up, integration readiness, and operational needs to support software across stack layers. What You'll Be Doing: Lead end-to-end execution of software platform initiatives, including planning, execution, delivery, and operationalization. Work together with software, infrastructure, product, and operations teams to ensure alignment on goals, deliverables, achievements, and schedules. Lead cross-functional initiatives encompassing cloud-native services, platform software, system integration, and release delivery. Help connect software roadmap execution to full-stack readiness, including dependencies across infrastructure, bring-up, validation, and downstream operational support. Identify cross-functional dependencies, mitigate risks, and drive resolution of complex technical and programmatic issues. Establish clear success metrics and reporting mechanis

KubernetesLinuxMachine LearningArtificial Intelligence
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

We’re hiring a Sr. Customer Success Manager to turn new deployments into measurable outcomes for engineering teams using Coder. You guide customers from onboarding through renewal, clearing blockers, and aligning our platform to their goals. You lead rollouts, monitor health and usage, and keep stakeholders informed. You partner with Sales on expansions and renewals, and with Support and Product on issues and feedback. What you’ll do here Own onboarding and rollout plans for new customers; coordinate with Sales for a smooth, successful implementation of Coder’s platform Remove barriers to adoption so customers achieve their desired outcomes expediently Engage with customers to understand goals, challenges, and use cases, and provide tailored guidance and recommendations Identify expansion opportunities within existing accounts, and partner with Sales on upsell and cross‑sell motions Oversee renewals and forecasting, ensuring timely, successful commitments Serve as the primary point of contact for inquiries, issues, and escalations, and collaborate with Coder teams to resolve Monitor and report on customer health and usage Develop deep expertise in Coder’s products to provide global customer coverage What we’re looking for 3+ years in software/SaaS sales and/or customer success, with enterprise experience preferred Working knowledge of cloud infrastructure, DevOps, developer tools, coding agents, platform engineering, CI/CD, and the SDLC Hands-on experience with Salesforce and other industry-standard CS platforms History of building strong relationships across executive, business, and technical stakeholders Consistent internal advocacy for customers and the ability to provide actionable feedback for Product and cross‑functional teams Habit of staying current on industry trends, CS best practices, and the competitive landscape Startup experience High EQ with strong verbal communication and technical writing skills Self‑motivated with a creative and analytical approach

AWSCI/CDAIGo
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -100%

From $200K/yr

Quick readStrong listing-quality and freshness signals

About the role Midjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species. We are a small, self-funded team focused on design, human infrastructure, and AI. We're looking for a senior product designer who's spent years building software. You'll design core product experiences across web and mobile, working directly with engineers and researchers from the earliest stages of a project through launch. What you'll do Own major product surfaces end-to-end: concept, flows, interaction, visual polish, ship Invent interaction paradigms for things that haven't existed before and make those new paradigms legible and fun Use research, experimentation, analytics, and user feedback to understand what people need and where products can improve. Prototype constantly, in Figma and in code — tools like Claude and Cursor should be part of your workflow Partner directly with engineering and research to refine interactions, iterate quickly, and ship high-quality work. You might be a fit if You've spent the last decade designing software people love to use. You're comfortable moving between interaction design, visual design, and product thinking. You enjoy working closely with engineers. You agree that good work often emerges out of many rounds of iteration. You explore wildly, are always ready to kill your darlings, and rework ideas until something clicks and feels correct. You'd rather prototype than make a presentation. Work > ego Logistics Full-time. Ideally San Francisco or New York, but PST/EST hours ok. Compensation: $200,000–$285,000 + benefits To apply A portfolio is required. We care much more about what you've made than where you've worked; if this sounds like you, apply even if you don't check every box.

D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%
Quick readStrong listing-quality and freshness signals

At Datadog, we’re on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale, enabling seamless collaboration and problem-solving among Dev, Ops, and Security teams globally for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Observability Data Platform (ODP) is the backbone of everything Datadog delivers – powering how data is ingested, stored, routed, and surfaced across every product at planet scale. As a Senior Product Manager for ODP, you will work with world-class engineers and cross-functional partners to shape how the platform is deployed, controlled, and operated. You will define product direction across the control plane and data layer, translate complex infrastructure trade-offs into clear roadmap decisions, and help customers get the most from their observability investment – regardless of architecture, topology, or scale. At Datadog, we place value in our office culture – the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You Will Do: Develop a deep understanding of the Observability Data Platform customers – platform engineers, SREs, and product managers that own the product verticals – their infrastructure challenges, deployment topologies, and cost-to-serve trade-offs. Define product direction across multiple ODP surfaces, including the control plane and data layer, by articulating clear problem statements and desired outcomes, and partnering with engineering on technical approach and sequencing Lead conversations with design partners and strategic customers to understand real-world platform pain points, validate product assumptions, and guide solutions from early prototypes through General Availability Develop a co

AIGoRustSpring
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%
Quick readStrong listing-quality and freshness signals

Datadog is looking for a Senior Product Manager to help lead the evolution of our fleet and lifecycle management capability, the product surface that gives customers visibility into, and control over, the observability software running across their infrastructure. This capability manages the deployment lifecycle for core observability agents and OpenTelemetry collectors running on customer hosts and containers. The Senior PM will expand the scope of fleet capability to additional Datadog software components, making it the single place customers go to see everything running in their environment, at any version, in any deployment model, and to manage it remotely and safely at scale, for both human operators and, increasingly, AI agents acting on their behalf. This is a high-visibility, cross-functional role. You'll partner with multiple engineering teams and be responsible for defining and delivering a coherent, unified fleet experience across UI, API, and MCP for customers. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do Own and evolve the product vision and roadmap for a unified fleet and lifecycle management capability spanning multiple product lines and deployment models. Define what "managed" means for each new software component as it's brought into fleet, balancing consistency of experience with the realities of each component's operational model. Drive a phased expansion plan, sequencing new components into fleet based on customer value, technical complexity, and dependency readiness. Partner closely with engineering leads across several teams to align on shared architecture principles to support disparate software components. Represent the voice of the customer for a capability that must work equally well for human operators using a UI and for AI

KubernetesAIGoRust
C
📍 Tampa Florida United States, United States
✓ Quality checkedCompany trend +800%

The Engineering Lead Analyst – Test Automation Platform Engineering is a senior-level technical leadership role responsible for driving the architecture, implementation, management, operational support, and continuous enhancement of enterprise test management and automation platforms. In this role, you will lead efforts to modernize test automation capabilities across the global technology ecosystem. You will architect end-to-end integration workflows, embed automated quality gates into enterprise CI/CD pipelines, and administer as well as operationally support both vendor and internally developed enterprise platforms (e.g., Core Performance Engineering / Performance Center, ALM-Quality Center, Zephyr Enterprise, CSDP / Octane). Additionally, you will play a critical role in production operations—delivering tier-3 platform support to rapidly and safely troubleshoot, triage, and remediate performance and availability issues in complex, distributed production environments. The ideal candidate blends deep hands-on expertise in software testing frameworks, modern DevOps pipelines, containerized infrastructure, message-driven integration, and robust operational resilience practices with strong governance, compliance, and stakeholder leadership skills. Key Responsibilities 1. Platform Engineering & Operational Support Install, configure, upgrade, administer, and support enterprise test management and performance engineering toolsets (e.g., Core Performance Engineering / Performance Center, ALM-Quality Center, Zephyr Enterprise, Core Software Development Platform [CSDP] / Octane). Provide end-to-end operational support for

DockerKubernetesArtificial IntelligenceAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

PythonCI/CDLinuxAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -84.1%
Quick readStrong listing-quality and freshness signals

About the Team The Emerging Products team is a lean, high-output product lab group that builds products at the forefront of model capabilities. We collaborate across all teams within the company, from research and infrastructure to consumer products. The team is responsible for identifying new product opportunities, building them quickly, dogfooding them internally, and then launching the successful products to users. We use data, user research, and analytics to inform our ideas, and make decisions on what experiments are worth iterating, stopping, or scaling. About the Role We’re looking for a senior, product-minded software engineer to own ambiguous 0-to-1 work from idea through prototype, validation, and handoff. This is a full-stack role with a strong frontend and product emphasis: you will build the interfaces and supporting backend systems needed to test new experiences quickly, while making sound architectural choices that enable successful concepts to scale. This role is based in our Mission Bay office in San Francisco. In this role, you will: Build and ship high-quality, product experiments across the full stack. Turn ambiguous user needs and emerging technical capabilities into testable product concepts, using research and metrics to guide iteration. Own technical direction for 0-to-1 projects, balancing speed, reliability, and a clear path from prototype to scalable product. Partner closely with design, product, research, and engineering teams to dogfood, evaluate, launch, and transition successful experiments. You might thrive in this role if you: Have a track record of building and shipping end-to-end products in fast-moving, startup, founder-led, growth, or other high-ownership environments. Bring strong frontend engineering skills and enough backend and systems depth to make sound full-stack architectural decisions. Pair product intuition with evidence, using user research and product data to identify opportunities and make pragmatic tradeoffs. Operat

AWSRestAIRust
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re hiring a People Business Partner to support our team during a critical phase of growth. This is a highly strategic, high-impact role for someone who has partnered closely with leadership teams and helped organizations scale with intention. You will be deeply embedded with managers, bringing clarity and rigor to how teams are structured, how leaders operate, and how talent is developed across the organization. You’ll help shape team effectiveness, identify critical talent gaps, drive talent and performance strategies that enable high-performing teams, and build the people practices and change management approaches that allow us to scale with both speed and discipline. This role requires strong business judgment and the ability to operate with deep context. You’ll partner closely with leaders to navigate complex organizational decisions, anticipate challenges before they surface, and bring a clear point of view on what great looks like at every level of the organization. RESPONSIBILITIES Strategic partnership to leadership Serve as the trusted people partner to leadership, maintaining deep business context and translating it into people priorities by proactively surfacing systemic issues, risks, and opportunities before they become urgent. Bring data-driven insights to advise management on org design, succession planning, performance, retention, and engagement. Build management capacity across the org, equ

Machine LearningAIGoRust
R
📍 Foster City, California, United States· Full-time
✓ Quality checkedCompany trend -87.5%

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: As a Staff Product Engineer at Replit, you’ll work closely with other product and platform engineers, designers, sales representative, and product managers to build features that help users collaborate with their team to go from idea to software fast. You’ll be at the forefront of shaping and experimenting on what our tens of millions of users love. You will: Help lead major projects and take new products from 0->1 Identify the hardest technical and/or quality problems holding us back, and then build solutions Chart high level technical direction and follow up to make sure those projects come together to deliver on results Mentor and develop new senior engineers to help grow the team Ship new features and build infrastructure using: TypeScript, React, CSS, GraphQL, Node.js, and Postgres Required skills and experience: A minimum of 7 years of professional software development experience Experience in a technical leadership role, working cross functionally Working experience building full stack applications with TypeScript Working experience building directly for users Bonus Points : You’re excited about the future of programming and have experience working with IDEs, terminals, or other common developer tools You’ve had previous experience working at a startup in a cross-functional engineering role This is a full-time role that can be held from our Foster City, CA office. The hybrid role has an in-office requirement of Monday, Wednesday, and Friday. Full-Time Employee Benefits Include: 💰 Competitive Salary & Equity 💹 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance 🩼 Short Term and Long Term Disability 🚼 Paid Parental, Medical, Caregiver Leave 🏝 Flexible

TypeScriptReactNode.jsGraphql
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As the Senior Product Manager for the Actions & Automations team, you will own the ecosystem that enables customers, partners, and Datadog teams to build, deploy, and operate AI agents on Datadog. You will drive the strategy and execution for the platform capabilities, developer experience, integrations, and extensibility model that make Datadog the best place to build agents that understand and act on production systems. Modern engineering organizations are entering a new era where software is not only monitored and operated by humans, but increasingly by AI-powered agents. As agentic workflows reshape how teams build, operate, secure, and troubleshoot systems, customers need a platform for creating specialized agents, connecting them to business and engineering systems, governing their behavior, and extending them to solve unique organizational problems. You will define and build the ecosystem that makes this possible. Agent Builder sits at the intersection of Datadog's products, AI capabilities, and ecosystem strategy. You will have the opportunity to work across the breadth of the Datadog platform, partner with teams throughout the company, and help establish Datadog as the foundation for operational AI. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define the vision, strategy, and roadmap for Datadog's Agent Builder platform and ecosystem. Own the core platform capabilities that enable customers and partners to create, customize, deploy, and manage AI agents. Drive the extensibility model for agents, including integrations, tools, actions, context sources, APIs, SDKs, and developer workflows. Shape how agents perform actions across Datadog products and third-party systems. Partner closely with AI, platform, infrastructure, and product tea

AIGoRustSpring
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As a Senior Platform Product Manager focused on AI SDLC Trusted Throughput, you will define and drive the product strategy for enabling safe, reliable software delivery at AI-native scale across Datadog’s Internal Developer Platform. As AI accelerates development velocity and system complexity, you will help evolve SDLC systems from human-supervised workflows to platforms with built-in safety, observability, and correctness guarantees. You will partner closely with engineering, security, and developer platform teams to improve deployment reliability, operational visibility, and governance while enabling both engineers and AI agents to move quickly with confidence. This role offers the opportunity to shape foundational developer infrastructure and influence how AI-powered software delivery operates across Datadog. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the product strategy, roadmap, and execution for AI-native SDLC throughput and reliability initiatives across Datadog’s Internal Developer Platform Define and drive platform outcomes aligned to DORA metrics, balancing deployment velocity with reliability, change failure reduction, and operational safety Partner with engineering, infrastructure, security, and developer experience teams to build automated validation, auditability, and risk-scoring capabilities into deployment workflows Deliver actionable SDLC observability and diagnostic capabilities that connect executive-level metrics to operational signals across the software delivery lifecycle Drive systems that monitor and validate AI-generated or AI-attributed changes to ensure correctness, compliance, and trustworthy automation Serve as a cross-functional product leader across SDLC Foundations, Security Engineering, and compl

AIGoRustSpring
C
📍 Tampa Florida United States, United States
✓ Quality checkedCompany trend +800%

The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=

JavaScriptTypeScriptPythonJava
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -83%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We are bringing our people systems in-house on Workday, and we are hiring our first dedicated Workday Analyst to make it excellent. You will join while the implementation is underway, ramp alongside our deployment partner, and own the tenant from go-live onward. This is a hands-on configuration role: you will build business processes, manage security, load data, and test releases yourself, and you will teach others on the People team to do the same as we grow. If you want to shape a Workday environment from its first day in production instead of inheriting years of someone else's decisions, this is that rare opening. RESPONSIBILITIES Own day-to-day Workday configuration: business processes, security groups and roles, custom reports, calculated fields, and tenant settings Build and run EIB loads for data changes, mass updates, and audits Own the twice-yearly Workday release cycle: evaluate new features, regression-test, and roll out changes safely Shadow the implementation build, then take over tenant ownership at go-live Monitor integrations and triage issues with our IT team and vendors Support payroll configuration in partnership with our Accounting team and payroll services provider Field and resolve system requests from employees, managers, and the People team Mentor teammates so Workday administration becomes a team capability, not a single point of failure REQUIREMENTS 4+ years administering a live Workday

Machine LearningAIAccountingPayroll
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl

PythonLinuxAIC++
🔔

Get new senior infrastructure engineer jobs in United States by email

Daily job updates · Unsubscribe anytime