We at NVIDIA seek an Senior Developer Relations, Automated Synthetic Chemistry Science Lead who can coordinate science streams. This role exists at the crossroads of experimental science, optimization, analytical characterization, AI modeling, laboratory automation, and data infrastructure. It demands a practitioner with a background in research who can lead interdisciplinary science teams and accelerate experimentation without compromising scientific quality. What you'll be doing: Coordinate the operating plan across research priorities, technical execution, and program achievements. Convert research objectives into experimental priorities, parameter-space development, campaign planning, and success criteria. Lead research initiatives encompassing synthetic chemistry, catalysis, process chemistry, machine learning, automated systems, and data processing. Build standardized experimental traces capturing successful and unsuccessful outcomes, metadata, quality-control signals, analytical summaries. Guide AI systems for feasibility assessment, outcome prediction, optimization, scope exploration, and campaign orchestration. Integrate automated experimentation, analytical data streams, data curation, model retraining, and campaign decisions into a closed-loop operating model. Establish science-stream governance: build reviews, decision logs, risk and dependency tracking, quality thresholds, and paths for addressing blocking issues. Prepare recurring workstream readouts for program leadership, including scientific progress, critical decisions, cross-team dependencies, resource needs, and unresolved risks. What we need to see: PhD (or equivalent experience) with 10+ years of hands-on experience in experimental science, chemical engineering, material science, robotics, a
Jobs in United States
Senior Infrastructure Automation Engineer in United States
2,153 active opportunities · Updated October 2026
Showing
15 jobs
Explore current senior infrastructure automation engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debug compl
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. You will build the operating system for EPD — the systems, agents, and practices that make engineering, product, and design unreasonably effective at delivery, with Jira as the substrate. The EPD Systems team is building the infrastructure Vanta's engineering, product, and design organization depends on to understand itself and operate well. We're constructing three interlocking layers: sources of truth at the foundation, a shared interpretive layer that reasons across all of them, and the operating system that makes the work itself legible. This role owns the operating system. This is not a traditional PMO or status-reporting role. You build — agents, automation, workflow design, and the practices that make teams want to use the system rather than route around it. What you’ll do as a Senior Systems Designer at Vanta: Design and own workflow and hierarchy across engineering, product, and design in Jira Make sure Jira reflects real practice, and real practice reflects what needs to show up in Jira — in both directions Build the agents and automation that run the system yourself, with AI as part of how you work Drive adoption: go to the teams whose practice doesn't match the substrate today and change that — through conversation, well-built artifacts, or clear instruction that works without you in the room Understand delivery breakdowns at the mechanism level and build fixes that address the root cause, not the symptom Build alongside teammates who are growing into more technical work — raise their ceiling, not just your own output How to be successful in this role: You build and ship, recently, with AI as part of how you work. "
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As an OS / K8s Systems Engineer at Baseten, you’ll build the automation and systems that turn raw GPU hardware into production-ready compute. From provisioning to orchestration, you’ll own the software layer that makes our infrastructure reproducible, scalable, and reliable across data centers. This is a senior, hands-on role focused on building systems not operating them. You’ll work close to the metal designing OS images, building provisioning pipelines, and automating cluster bring-up from scratch. Your work will define how quickly we can turn new capacity into usable compute. EXAMPLE INITIATIVES Zero-to-cluster automation Build workflows that take new hardware from unprovisioned to fully operational cluster. Provisioning systems Design PXE-based or equivalent systems for imaging and lifecycle management. Reproducible infrastructure — Ensure clusters deploy consistently across data centers. RESPONSIBILITIES Own the end-to-end automation of cluster bring-up and lifecycle management. Build and maintain OS images, provisioning systems, and configuration pipelines. Deploy and operate cluster orchestration platforms (Kubernetes, Slurm, or similar). Design systems for reproducibility across sites and hardware generations. Automate upgrades, rollouts, and failure recovery. Optimize system performance, including GPU utilization and networking. Partner with hardware and network teams to validate and improve system b
From $192K/yr
As a Senior Platform Product Manager focused on AI SDLC Trusted Throughput, you will define and drive the product strategy for enabling safe, reliable software delivery at AI-native scale across Datadog’s Internal Developer Platform. As AI accelerates development velocity and system complexity, you will help evolve SDLC systems from human-supervised workflows to platforms with built-in safety, observability, and correctness guarantees. You will partner closely with engineering, security, and developer platform teams to improve deployment reliability, operational visibility, and governance while enabling both engineers and AI agents to move quickly with confidence. This role offers the opportunity to shape foundational developer infrastructure and influence how AI-powered software delivery operates across Datadog. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own the product strategy, roadmap, and execution for AI-native SDLC throughput and reliability initiatives across Datadog’s Internal Developer Platform Define and drive platform outcomes aligned to DORA metrics, balancing deployment velocity with reliability, change failure reduction, and operational safety Partner with engineering, infrastructure, security, and developer experience teams to build automated validation, auditability, and risk-scoring capabilities into deployment workflows Deliver actionable SDLC observability and diagnostic capabilities that connect executive-level metrics to operational signals across the software delivery lifecycle Drive systems that monitor and validate AI-generated or AI-attributed changes to ensure correctness, compliance, and trustworthy automation Serve as a cross-functional product leader across SDLC Foundations, Security Engineering, and compl
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track
Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu
From $154K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the world, delivering the object, block, and file storage platforms that power GoDaddy's hosting infrastructure, internal services, OpenStack environments, and next-generation AI/HPC workloads. If you're passionate about distributed systems, storage architecture, and solving failure scenarios at massive scale, this is an opportunity to work on infrastructure few engineers will experience in their careers. Ceph is a strategic platform at GoDaddy — not an ancillary service. Our global footprint includes 80+ production clusters, 20,000+ OSDs, 1,830 storage nodes, 300 PB of raw capacity, and 69 billion objects spanning five datacenters across three continents. The platform supports RBD, RGW (S3/Swift), and CephFS workloads through more than 1,550 pools, 574,000 placement groups, and 900+ MDS daemons, creating engineering challenges that demand deep expertise in storage architecture, data durability, performance optimization, automation, and observability. As a Lead Senior Site Reliability Engineer, you'll serve as one of the principal technical leaders for GoDaddy's Ceph platform. You'll design the next generation of storage clusters, lead major platform upgrades, drive capacity and hardware strategy, and establish the standards that govern how the platform scales. You'll be the engineer the team turns to for the most complex s
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h
$145K – $170K/yr
CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. CLEAR is seeking a Senior Security Operations Analyst III to join our SOC team to help strengthen our ability to detect, investigate, and respond to evolving security threats. In this role, you’ll lead complex investigations, improve CLEAR’s threat detection and response capabilities, and serve as a trusted security partner while helping develop the analysts and program around you. What you'll do: Lead complex investigations of security events across corporate networks, endpoints, data centers, cloud environments, and other critical systems, driving incidents from initial analysis through escalation and remediation Develop, tune, and optimize threat detection logic across SIEM, EDR, and other security platforms, proactively identifying coverage gaps, reducing false positives, and improving the fidelity of security alerts Partner with Engineering, Infrastructure, and other teams to investigate threats, identify root causes, communicate risk, and drive timely remediation and improvements to CLEAR’s security posture Apply threat intelligence, data, automation, and AI-enabled tools to identify emerging attack patterns, accelerate investigations, improve detection workflows, and strengthen decision-making while applying sound security judgment Serve as a subject matter expert and escalation point for other analysts, mentoring junior team members, sharing knowledge, and helping establish scalable processes, playbooks, and standards for threat detection and analysis Continuously evaluate CLEAR’s detection coverage against the evolving t
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. About The role As a Vulnerability Management Engineer, you will support the identification, assessment, prioritization, and remediation of vulnerabilities across applications and infrastructure. Working under the guidance of senior team members, you will assist in understanding how vulnerable dependencies enter an application, identifying remediation options, and engaging with engineering teams to track fixes. You will contribute to the maintenance of internal vulnerability-management tools, such as scripts, documentation, and reporting. The ideal candidate will have a desire to grow their AppSec expertise, will be eager to learn about modern security tooling and automation, and will be comfortable using AI tools like Claude to assist with documentation, investigation, and scripting tasks while following company security and data-handling requirements. What you’ll do Perform regular vulnerability assessments using different tools. Regularly drive remediation and reporting of cataloged vulnerabilities. Assess discovered vulnerabilities and properly prioritize their scope, impact and necessary response actions. Conduct security reviews of our products and production infrastructure. Contribute to vulnerability management, application security and/or offensive/red-team operations. Engage in security audit
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. As a Senior Data Analyst on the Finance team, you’ll play a pivotal role in uncovering the insights that shape company-wide planning, forecasting, and strategic decision-making. This high-impact role blends analytical rigor, technical fluency, and business acumen to steer strategy and identify opportunities for scalable growth. The Finance Data team partners with both Finance and other teams across Vanta. You’ll collaborate closely with stakeholders to build trust and deliver data-driven recommendations; your work will influence everything from operating plans and investment decisions to self-serve reporting and process automation. This is a unique opportunity to deepen your analytics expertise in a dynamic, high-growth environment, with visibility into executive-level priorities and the ability to shape them with data. What you’ll do as a Senior Data Analyst at Vanta: Deliver strategic insights by conducting deep, prioritized analyses that inform high-impact financial decisions. Act as a technical partner to teams across FP&A, Accounting, RevOps, Product, Marketing, and more. Increase team efficiency by automating recurring deliverables and helping scale workflows as the company grows. Empower decision-makers with self-service dashboards and data assets that surface trends and drive informed choices across the business. Advance our analytics infrastructure by partnering with Data Engineering and other analytics teams to democratize access to high quality data and insights. How to be successful in this role: Experience: 4+ years of experience in data analysis or equivalent function. Exposure to FP&A, Accounting, or SaaS
From $128K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor
About the Team The Privacy Engineering team builds the systems and technical foundations that govern how user data is understood, retained, accessed, and used across OpenAI. We partner with Product, Data, Infrastructure, Security, and Legal to translate policy and trust commitments into durable architecture and enforceable controls. Our work spans data inventory and mapping, classification and lineage, retention and deletion, access governance, purpose and usage controls, auditability, and lifecycle automation. We aim to make policy-aligned data handling the default while giving teams clear, reliable primitives for building and operating products at scale. About the Role We are looking for an experienced Software Engineer to drive the architecture and execution of user data governance across OpenAI. You will define technical direction, build shared platforms and controls, and lead cross-functional programs that make data flows discoverable, policies enforceable, and ownership explicit. This role is well suited to a senior engineer who can move between deep systems design and organization-wide influence, turn ambiguous requirements into pragmatic roadmaps, and operate high-trust systems end to end. This position is based in San Francisco. Relocation assistance is available. In this role, you will: Set the technical strategy and architecture for user data governance across data mapping, classification, lineage, retention, deletion, access, and permitted usage. Design and build shared services, APIs, metadata systems, and policy-enforcement mechanisms that make governance controls consistent, scalable, and auditable. Establish reliable inventories of user data, system ownership, data flows, and policy applicability across products, infrastructure, analytics, and research systems. Partner with Product, Data, Infrastructure, Security, and Legal leaders to define decision rights, translate requirements into controls, and drive adoption across teams. Own governance systems
Cloud Platform Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations is currently seeking a Cloud Platform Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking a skilled platform engineer to help build and operate the cloud platform services that host critical enterprise applications and software toolchains. In this role, the selected candidate will focus on the shared platform capabilities that enable teams to deploy, run, and maintain containerized and cloud-hosted solutions in a consistent and supportable manner. As both an individual contributor and technical leader, this position will help define and implement platform standards for Kubernetes, container hosting, deployment automation, configuration management, and operational support. This role is focused on platform reliability, repeatability, scalability, and service enablement, rather than custom application software development. Position Responsibilities: Design, implement, and maintain cloud platform services supporting Kubernetes, containers, ingress, storage integration, secrets management, and service connectivity Build and sustain reusable deployment patterns for Commercial-Off-The-Shelf (COTS), Open Source Software (OSS), and internally customized applications Develop and maintain automation for platform provisioning, upgrades, patching, and lifecycle support Manage cluster lifecycle activities including: Cluster upgrades Node management <
Other cities to consider
More places hiring for this role
Get new senior infrastructure automation engineer jobs in United States by email
Daily job updates · Unsubscribe anytime