Build the infrastructure that keeps every NVIDIA chip aligned from first spec to final shipment. NVIDIA's Silicon Co-Design Group sits at the convergence of architecture, silicon, systems, and manufacturing. The System–Manufacturing Architecture (SMAC) team coordinates between system specifications and manufacturing test specifications from pre-silicon POR through production release across GPU, SoC, and CPU programs. When that alignment drifts, silicon faces the consequences: escapes, yield loss, and performance loss. We're hiring a Senior Manufacturing & System Co-Design Workflow Engineer to lead the methodology and infrastructure that maintains holistic, systematic alignment, at scale across the full portfolio. The strongest candidates in this role design the workflow before being asked to fix a program, and build the checks and automation that confirm alignment holds long after they've moved on to the next problem. What you’ll be doing: SMAC Workflow Methodology: Define manufacturing spec types, including schema and semantics, derived from system PORs and features. Own the methodology that governs how specification work gets structured, versioned, and validated across the program lifecycle. Production Python Pipelines & Automated Checks: Develop production-grade Python pipelines and automated checks that catch specification drift between system POR and manufacturing test programs ,ATE, SLT, BLT, L10+, before silicon exposes the discrepancy. The goal is that misalignments surface in the workflow, not on the tester. E2E Program Integration & TPM Attestation: Wire SMAC work into the end-to-end program spine, milestones, gates, and artifacts, and define explicit TPM-driven attestation when checks lag. Alignment can't be assumed; it must be proven at every stage. Agent-Ready Tooling & CI Infrastructure: Integrate tooling into an agent-ready
Jobs in United States
Senior Infrastructure Automation Engineer in United States
2,189 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior infrastructure automation engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and dedicated people on the planet working for us. If you're creative and autonomous, we want to hear from you! What you'll be doing: Developing software to enable efficient network design, deployment and day 2 management. Building product focused software solutions, used by internal and external customers. Helping us as we transform our workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies. Owning and driving integrations with various service APIs such as Cloud Service Providers, to automate creation of environments and auto populate data sources in turn. Building on open source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations. Streamlining deployment mechanisms and life cycle operations Developing modern service architectures around streaming data and event pipelines. Working with infrastructure domain experts on true, zero touch deployment solutions and utilizing best of breed high performance computing management solutions. Be a proactive problem solver, looking out for new opportunities to improve our services and customer experience. Communicate readily with your peers across the organization, b
From $10K/yr
About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role You'll be the architect of our endpoint security posture across the full fleet — macOS, Windows, and BYOD mobile devices. You'll build secure-by-default controls with Terraform and GitOps, harden endpoints at scale, and automate the full device lifecycle so nothing depends on a human remembering to click the right button. You'll partner with IT, SecOps, and engineering teams to sharpen our telemetry and detections, mentor other engineers, and raise the bar on what "low-friction security" actually means. Everything you ship will be auditable, measurable, and built for the long run. We're also thinking seriously about how corporate security evolves in an agentic world — where AI agents act on behalf of employees and traditional identity and endpoint assumptions break down. You'll help us shape that answer. What You’ll Do Write and ship MDM policy as code — configuration profiles, remediation scripts, and enforcement rules across macOS, Windows, and mobile — with staged rollouts and rollback from day one Build patch automation that close
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will contribute to the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with hardware, software, validation, and manufacturing teams, you will develop test content, automation, diagnostics, and characterization capabilities that support silicon bring-up, yield learning, and manufacturing deployment. The ideal candidate will have strong technical foundations in semiconductor test and validation, excellent debug skills, and a passion for improving product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Develop and maintain SLT test content, automation, d
We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment. What you'll be doing: Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency Leverage AI skills to expedite the test scope, test plan, execution and automation workflows. Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas. Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed. Perform performance, scalability, and reliability testing of cloud services. Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud. Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability. Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues. Working with tea
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe
NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production. Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. What you'll be doing: Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation. Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads. Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads. Devel
From $195K/yr
Here at Datadog, we think about offensive security a little bit differently. We embrace automation and AI to run adversary simulations continuously across a massive cloud-native environment, and we expect our offensive engineers to build the tooling that makes that possible. We're looking for a Senior Security Engineer who can execute sophisticated red team operations, write the code that scales them, and take an AI-first approach to offensive security engineering. At Datadog, we place value in our office culture - the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You'll Do: Plan and execute red team engagements end-to-end, simulating real-world threat actors across cloud infrastructure (AWS, GCP), Kubernetes, CI/CD pipelines, and corporate environments Build and maintain custom offensive tooling, automation frameworks, and engagement infrastructure, treating offensive operations as a software engineering problem Develop custom payloads and evasion capabilities tailored to Datadog's environment and modern defensive controls (EDR, SIEM, network monitoring) Improve the efficiency of offensive operations through thoughtful use of automation and AI, accelerating reconnaissance, vulnerability analysis, and reporting workflows Partner with the Detection & Response team on purple team exercises to validate detection logic, improve alert fidelity, and influence threat models Translate offensive findings into concrete improvements by working directly with defensive security and engineering teams to close gaps Who You Are: You have 5+ years of hands-on experience in offensive security (red teaming, penetration testing, or adversary simulation) with a track record of operating against mature, well-defended environments You write production-quality code (Python, Go, or similar), can build your own tools, and automate your w
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Software Engineers deliver high-value products for our customers, and infrastructure that enables our business to scale. Vanta’s team and technology surface are growing quickly, and it’s essential that we invest in the right abstractions and systems to enable us to scale with our business. As a Software Engineer, you’ll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, and driving complex projects across our technical stack.. Your past experience will be leveraged to enable and accelerate Vanta’s growth. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. Our product integrates deeply with the services that present security risk to a company, pulls and analyzes data from those sources, and surfaces potential security threats to our customers in real time with guidance to remediate them. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We’re growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. As a Software Engineer on the Expansion Team, you’ll play a central role in shaping Vanta’s self-serve and product-led growth motion. Instead of working horizontally across compliance workflows, this team focuses on the full post-activation lifecycle—driving retention, expansion, and revenue automation across the product. You’ll build the core systems that power Vanta’s purchasi
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. As a Senior Software Engineer, you'll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, driving complex projects across our technical stack, and mentoring our talented engineering team. Your past experience will be leveraged to enable and accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Authoring Experience team owns how Vanta tests are authored and deployed to customers, from the APIs and services underneath to the experiences built on top. As a Senior Backend Engineer on this team, you will own the services and public-facing customer APIs that turn Vanta's automation platform into products customers touch directly, stitching together systems across the platform to serve them well. What you’ll do as a Senior Fullstack Software Engineer at Vanta: Own the architecture and rollout strategy for critical platform services, from initial design through large-scale rollout and migration. Set the technical bar for how we design, publish, and maintain customer-facing APIs while owning things like security, reliability, and developer experience. Stitch together
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. Vanta’s team and technology surface are growing quickly, and it’s essential that we invest in the right abstractions and systems to enable us to scale with our business. As a Senior Software Engineer, you’ll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, driving complex projects across our technical stack, and mentoring our talented engineering team. Your past experience will be leveraged to enable and accelerate Vanta’s growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We’re growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Integrations Platform team mission is to power the world’s largest trust automation ecosystem, enabling any person or agent to build, connect, and automate trust seamlessly. We own Vanta’s integration ecosystem, which currently includes over 400 integrations across Cloud Providers (AWS, Azure, GCP), Identity Providers, Mobile Device Management (MDM), and Human Resources Information System (HRIS). We are focused on developing the Integration Platform. This includes creating shared primitives for authentication, lifecycle, observability, and publishing to ensure all integrations are built on the same fou
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Software Engineers build and own high-value products for our customers and the infrastructure that lets our business scale. Vanta's team and technology surface are growing quickly, and it's essential that we invest in the right abstractions and systems to scale with our business. As a Software Engineer, you'll build and own full-stack features and platform primitives across our stack, work closely with the teams and partners who depend on what you ship, and grow into broader technical ownership. Your work will directly accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Integrations Platform team mission is to power the world's largest trust automation ecosystem, enabling any person or agent to build, connect, and automate trust seamlessly. We own Vanta's integration ecosystem, which currently includes over 400 integrations across Cloud Providers (AWS, Azure, GCP), Identity Providers, Mobile Device Management (MDM), and Human Resources Information System (HRIS). We are focused on developing the Integration Platform. This includes creating shared primitives for authentication, lifecycle, observability, and publishing to ensure all integrations are built on the same foundation. Our North Star is to eliminate the barrier to building integrations entirely. We aim to enable any
From $243.3K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service discovery, secrets management and related software layers. We’re looking for a skilled Senior Site Reliability Engineer with strong programming skills to help us build Roblox's private cloud, productionize our growing Kubernetes-based infrastructure, and institute reliability best practices across the Roblox Compute team. You will: Design and Develop systems & libraries that promote fault-tolerance and resilience, automate much of the management and lifecycle of our clusters, and ensure systems are observable. Promote and Institute reliability best practices across the Infra Compute group, drive common reliability initiatives. Provides collaborative technical reviews and operational guidance to strengthen system reliability. Build, Automate and Standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem. Create Tooling that provides production guardrails, by evaluating release candidate capacity with load testing tooling before de
About Datadog We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at a high scale - trillions of data points per day — providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way. The Opportunity We are looking for an experienced software engineer to join our CI/CD Security Team within our SDLC Security organization. We work at the intersection of security and engineering infrastructure to secure Datadog's continuous integration and continuous delivery systems. Our responsibilities include hardening pipelines, protecting credentials, and enforcing tightly scoped access controls. We also develop authorization and verification mechanisms to ensure that only trusted code and approved processes can reach production. In this role, you will shape and build a new security layer for our CI/CD infrastructure and drive its adoption across the engineering organization. You will solve challenging systems problems around trusted build provenance, secure secret delivery, and real-time policy enforcement at high throughput. The work sits directly in the critical path of software delivery, where strong security guarantees have to coexist with low latency, high reliability, and a seamless developer experience. You’ll join at an ideal time to make a big impact, as the need for robust software supply chain security is higher than ever. Datadog is growing rapidly, and AI-assisted development is increasing both the pace of software delivery and the amount of activity flowing through our CI/CD systems. Securing that scale without slowing engineers down requires strong software engineering fundamentals, thoughtful automation, and security controls designed to operate reliably at high throughput. At Datadog, we pla
$192K – $240K/yr
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corporate cards and banking with intuitive spend management, bill pay, and travel software, Brex enables founders and finance teams to accelerate operations, gain real-time visibility, and control spend effortlessly. Brex’s AI-native automation and world-class service eliminate manual expense and accounting tasks for customers so they can focus on what matters most. Tens of thousands of the world's best companies run on Brex, including DoorDash, Coinbase, Robinhood, Zoom, Plaid, Reddit, and SeatGeek. Working at Brex allows you to push your limits, challenge the status quo, and collaborate with some of the brightest minds in the industry. We’re committed to building a diverse team and inclusive culture and believe your potential should only be limited by how big you can dream. We make this a reality by empowering you with the tools, resources, and support you need to grow your career. Engineering at Brex Engineering at Brex is about building systems that scale with speed and intention. Our teams span Software, Data, Security, and IT, and operate with high autonomy and deep collaboration. We tackle hard technical problems, own our outcomes, and push for excellence at every level — from architecture to deployment. It’s an environment where engineering is a craft, and builders become leaders. What you’ll do As a Senior Software Engineer, Infrastructure (Release Engineering) at Brex, you will design, build, and operate the core systems that power Brex’s release, observability, and incident management processes. You will partner closely with product, platform, and operations teams to ensure releases are safe, fast, and reliable, and that our infrastructure scales securely as Brex grows. Where you’ll work This role will be based in our San Francisco office. We are a hybrid environment that combines the energy and connect
Other cities to consider
More places hiring for this role
Get new senior infrastructure automation engineer jobs in United States by email
Daily job updates · Unsubscribe anytime