Jobs in United States

Senior Infrastructure Automation Engineer in United States

2,153 active opportunities · Updated October 2026

Explore current senior infrastructure automation engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Why Safety? At Roblox, we strive to connect a billion people with optimism and civility, and the Safety organization’s mission is to become the leader in civil immersive online communities. We systematically detect, remove, and prevent problematic content and behavior, and we make Roblox accounts secure and free from compromise. We cover a broad area of the tech spectrum, including machine learning, classifiers for 3D models, experimentation, automation, detection workflows, and AI-powered text filters. Aligned and partnering with product teams, we use this tool belt to discover new opportunities, influence and shape the product roadmap and prioritization, build safety products, and measure the impact on our community of users and developers. In doing so, we keep Roblox safe, civil, and inclusive, and we foster positive relationships between people around the world. Why Safety Foundation? Safety Foundation is the infrastructure backbone that powers safety across Roblox. Every piece of content reviewed, every moderation action taken, every safety workflow running anywhere on the platform flows through systems built by this team. We are the platform that other safety teams build on — and righ

SQLAWSGitRest
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $187.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior QA Engineer, you will be the first dedicated QA hire for the Safety Engineering Group, which is responsible for building the infrastructure and policies that make Roblox the safest online community in the world. You will report to the QA Engineering Manager for Universal Apps and collaborate across the Safety organization to own end-to-end test strategies for high-stakes, 24/7 incident response systems and parent-facing safety features. This role is a unique opportunity to shape the quality culture of a company-level mandate from the ground up. You will: Partner with Engineering and Product to define and maintain comprehensive, risk-based test coverage for Safety features and systems. Own the end-to-end test strategy for Safety initiatives, including functional, regression, integration, and edge-case validation. Leverage and expand internal automation frameworks to automate critical user flows and drive API-first validation strategies. Identify high-impact automation opportunities and incorporate industry best practices to design robust, abuse-resistant test scenarios. Develop, track, and report on quality metrics (e.g., defect trends, coverage gaps) to proactively surface relea

AWSGitRestAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Software Engineer on our Release Engineering team, you will build the systems that hundreds of engineers across Roblox use every day to safely ship code to tens of millions of concurrent players. Your work will empower teams to ship bold, high-impact changes quickly and confidently on every device Roblox runs on. If you enjoy building developer-facing infrastructure where reliability and blast radius directly impact end-users, you will be right at home on our growing team. You Will: Design and develop backend services and automation that power our release and experimentation process across desktop, mobile, console, VR, and servers Work in C++ engine code to improve telemetry reporting, enhance feature rollout and automatic abort capabilities, and extend release functionality Build progressive client and server rollout, regression detection, and automated rollback systems that keep our weekly multi-platform releases safe at scale Work directly with engineering customers to turn pain points into durable, flexible, and safe-by-default tooling You Have: 5+ years building backend services with C#, Python, TypeScript, or similar Familiar with and comfortable working with C++ Familiar

TypeScriptPythonAWSGit
S
📍 Mclean, Virginia, United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer, Snowflake Natsec Running Snowflake in public sectors in different countries and regions, even in different industry verticals, requires us to build a compliant, secure, and auditable infrastructure. Many key design decisions are deeply rooted in the Snowflake product architecture. As a Senior Software Engineer, you will be responsible for leading several key areas and collaborating with various engineering groups in addition to the Public Sector team. To be successful in the area, you will need to have (and continue to build) a broad and in-depth knowledge base on cloud infrastructure, privacy, and governance, compliance controls, data security and data residency in various aspects of Snowflake. AS A SENIOR SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL: Solve real business needs at large scale by applying your software engineering and analytical problem solving skills. Design, implement and maintain scalable distributed systems for our cloud automation platform that include cloud control plane, Kubernetes container platform and traffic and networking. Work directly with customers to quickly understand their critical problems and design and implement solutions Deploy and maintain availability of cloud compute servers and Kubernetes cluster that power the

PythonJavaAWSAzure
MT
📍 Boise, ID - Main Site, United States
✓ Quality checkedCompany trend +1166.7%

Our vision is to transform how the world uses information to enrich life for all . Micron Technology is a world leader in innovating memory and storage solutions that accelerate the transformation of information into intelligence, inspiring the world to learn, communicate and advance faster than ever. As part of Micron's Technology Engineering & Innovation (TE&I) organization, you will have the opportunity to shape the future of our global infrastructure platforms while enabling business growth, operational resilience, and digital transformation at scale. The Opportunity Micron is seeking a transformational Senior Director of Technology Engineering & Innovation (TE&I) to lead the strategy, engineering, operations, and modernization of our global infrastructure ecosystem. This role is responsible for defining and executing Micron's vision across enterprise networks, cloud platforms, data center strategy, database services, infrastructure engineering, automation, observability, and global infrastructure operations. As a key member of the TE&I leadership team, you will drive innovation, operational excellence, and strategic transformation while building a high-performing organization focused on business outcomes and exceptional customer experiences. The successful candidate will be equally comfortable developing multi-year technology strategies, leading large-scale infrastructure transformations, driving operational performance, developing talent, and fostering a culture of collaboration, accountability, and continuous improvement. What You Will Do Lead Global Infrastructure Strategy & Transformation Define and execute Micron's global infrastructure vision and strategy. Develop multi-year roadmaps for: Enterprise Network Services <l

AIRecruitment
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -12.7%

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable &#34;flight recorder&#34; for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as &#34;Hardware Fault,&#34; &#34;Software Bug,&#34; or &#34;Environment Issue.&#34; This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define &#34;signals of impending failure,&#34; enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6&#43; years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
V
📍 United States· Full-time
✓ Quality checkedCompany trend -92.7%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Corporate Engineering team is the infrastructure layer that keeps 1,500+ people connected, secure, and moving fast. This role leads the shift from a reactive support function to a platform function the whole company relies on. Corporate Engineering owns the systems, tooling, and access infrastructure Vanta runs on. From identity and device management to AI tooling infrastructure and ITSM, CE makes sure Vanta's people have the right access, the right tools, and the right guardrails on day one and every day after. CE is at an inflection point. The company has scaled quickly, AI tooling is live company-wide, and the team is ready for a leader who sets technical direction, establishes clear ownership, and builds the leadership layer beneath them. As Sr. Manager, Systems Engineering, you will lead a team of experienced engineers, own several of the company's highest-priority platform initiatives, and make CE a function the rest of Vanta builds on rather than routes around. What you’ll do as a Senior Manager, Systems Engineering at Vanta: Lead, coach, and grow a team of systems engineers and a tech lead. Set clear expectations, hold a high bar, and develop the people who build Vanta's most critical internal systems. Own joiner, mover, and leaver automation end to end in Okta Workflows, so access is correct on day one, follows people through role changes, and closes completely on departure, including API keys and service credentials. Build and own device assurance across macOS and Windows, defining OS, browser, and device-state policy ahead of enforcement rather than in response to it. Rationalize the ITSM estate so self-servi

ReactGitRestAI
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -97.2%

From $127K/yr

Quick readStrong listing-quality and freshness signals

The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization. Among these are our multi-cloud-provider Kubernetes infrastructure, deployment machinery, and observability and alerting systems. The Fabric team manages the infrastructure that enables secure communication between systems and from the public internet. Their responsibilities encompass network architecture, service mesh, and edge load balancing, ensuring customer data remains safe in transit. The team plays a crucial role in developing and maintaining the reliable and globally connected multi-cloud network that supports MongoDB products. This role can sit in our NYC HQ, our smaller Austin, Palo Alto, or San Francisco offices, or fully remote from anywhere in North America. When based in an office, we provide hybrid work accommodation. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong networking background to join the Fabric team. This role is pivotal in building and maintaining the robust infrastructure necessary for secure and efficient communication between our services. As an SRE on the Fabric team, you will leverage your expertise in networking, distributed systems, and automation to ensure our systems are resilient, scalable, and reliable. The ideal candidate should Have 10+ years of experience working on software and operating distributed systems, with deep expertise in networking fundamentals and a good understanding of how the internet works, e.g. TCP/IP (including IPv6), DNS, TLS/mTLS, BGP, tunnels, overlays, and SDN principles Possess a customer-focused mindset, driving improvements that benefit end-users Value efficiency in processes and operations, and display a strong preference for automation over manual processes (“allergic to ops work”) Be intimately familiar with modern cloud-based infrastructure and the network design prim

MongoDBAWSAzureGCP
P
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. The Opportunity Postman is hiring a Senior Manager, Customer Success Engineering to help lead Customer Success Engineering in North America. This leader will manage CSEs, ensuring the team is focused on the customers where Postman can create outsized impact, while providing the coaching, inspection, and operational discipline required for consistently high-quality execution. This is a highly cross-functional leadership role and a critical partner to Sales. The leader will work closely with North America Sales leadership, with a strong focus on East Coast alignment, to prioritize accounts, shape technical engagement strategy, allocate CSE capacity, and ensure customer execution is progressing with urgency and discipline. CSEs help customers embed Postman into real engineering workflows across API design, governance, CI/CD, developer platforms, integrations, modernization, migration, automation, quality enforcement, and service discoverability. This is not about general support or enablement; CSEs are laser-focused on turning Postman from a beloved developer tool into critical API infrastructure. The right leader is an experienced po

ReactCI/CDAIGo
B
📍 Berkeley, United States
✓ Quality checkedCompany trend +7.9%

Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul

AzureTerraformAnsibleSap
R
📍 New York, NY, United States· Full-time· Remote
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp is in a critical phase of growth. We grew immensely last year and are building out a talented business systems team to ensure we maintain this trajectory for years to come. You’ll work directly with our Sales, Account Management, Partnerships, and Product teams to execute mission-critical business systems projects across the organization. This is a key role where you will be uniquely positioned to impact the full picture of Ramp’s growth efforts through systems development. What You’ll Do Work alongside Sales Operations to administer key go-to-market business systems, including Salesforce, Outreach, Qualified, Zendesk, Hubspot, Looker, Gong.io Build and deploy automation (flows), validations, and applications in Salesforce Implement new systems and integrations as needed Analyze key business requirements and systems capabilities to write specifications for systems build and run end to end implementation Create key reports and dashboards to track systems performance and data accuracy Write and maintain clear documentation on syste

RestAIGoSalesforce
H
📍 Boston, Massachusetts, United States· Full-time
✓ High-confidence listing

From $91.1K/yr

Quick readStrong listing-quality and freshness signals

We take play seriously. We’re looking for curious adventurers ready to find their party, fueled by imagination and drive to build what’s never been built before. At Hasbro and Wizards of the Coast, you’ll collaborate with passionate teams to reimagine our iconic brands and create experiences that spark joy, connection, and community through the magic of play. This is your chance to shape legendary play that lasts a lifetime. Reliable workplace technology enables teams to do their best work. This role focuses on maintaining and improving that reliability across devices, collaboration systems, and workplace technology environments. The Workplace Technology Support Engineer - Executive Support helps maintain the stability, performance, and continuous improvement of workplace technology across the organization, with a main focus on our executive leadership team. This will serve as the main point of contact for issues in the executive space. The role will need to be available to travel to support events, shows, presentations, and calls. You'll work closely with platform engineering, security, networking, and other internal teams to diagnose advanced issues, maintain operational standards, and translate recurring problems into scalable solutions through standardization, documentation, automation and AI-assisted operational improvements! Effective from the date that Hasbro opens its new Boston location, this position will be onsite Monday – Friday at Hasbro’s new HQ location in Boston, MA. In the interim, this position will be onsite Monday – Friday at Hasbro’s HQ in Pawtucket, RI. What You'll Do Executive Support Primary technology support resource for executives and senior leadership, delivering responsive, high-touch technical support across office, remote, and travel environments. Provide white-glove support for executive workstations, laptops, mobile devices, collaboration platforms, and workplace technologies while maintaining exceptional custom

GitAIGoSEM
P
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 CA, United States· Full-time· Remote
✓ High-confidence listingCompany trend -86.4%

From $1.7M/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . The Production Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams' capability to design, build and operate robust systems at scale. Pinterest's applications and infrastructure handle billions of monthly page views and petabytes of data as Pinterest continues to grow and scale. As a Senior Production Engineer on Solutions Engineering, you will design and build AI agents, platforms, tools, frameworks and methodologies to assure the reliability of our large-scale distributed systems serving hundreds of millions of monthly active users, handling hundreds of thousands of requests per second, and managing tens of petabytes of data. You'll lead infrastructure modernization initiatives, build intelligent automation that eliminates operational toil and amplifies engineer

PythonSQLMySQLAWS
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -83.9%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA

PythonSQLPostgreSQLMySQL
R
📍 New York, NY, United States· Full-time· Remote
✓ High-confidence listingCompany trend -99.2%

From $10K/yr

Quick readStrong listing-quality and freshness signals

About Ramp Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books. The problems are high-stakes, data-dense, and unforgiving. We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome. The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same. If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it. About the Role Ramp is seeking its first dedicated Financial Crimes Compliance Controls Senior Strategist to build and own the framework that demonstrates FCC controls are working as intended, remain effective as risk evolves, and can be clearly supported in bank-partner, audit, and regulatory examinations. This is a senior individual-contributor role at the intersection of financial-crimes compliance, controls effectiveness, technology, data, and AI-enabled operations. You will own FCC’s approach to testing, documenting, and improving manual and automated controls, including procedures, detection rules, AI and LLM-enabled workflows, and automation deployments. You will serve as FCC’s dedicated compliance partner for program-level Product, Engineering, Design, and Data (“PEDD”) initiatives, ensuring FCC requirements, controls, and risk considerations are incorporated into strategic product and technology planning. This role complements the broader team’s cross-functional partnerships by focusing specifically on major PEDD-led initiatives. This is a

RestMachine LearningAIGo
🔔

Get new senior infrastructure automation engineer jobs in United States by email

Daily job updates · Unsubscribe anytime