NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem supporting NVIDIA's GPU Cloud and NVIDIA SuperPod deployments. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and dedicated people on the planet working for us. If you're creative and autonomous, we want to hear from you! What you'll be doing: Developing software to enable efficient network design, deployment and day 2 management. Building product focused software solutions, used by internal and external customers. Helping us as we transform our workflows and organization into a centrally orchestrated configuration management framework, operating at scale across geographies. Owning and driving integrations with various service APIs such as Cloud Service Providers, to automate creation of environments and auto populate data sources in turn. Building on open source software, designing and implementing data structures and UI interfaces to automate processes from equipment purchase to device config generation to deployment to operations. Streamlining deployment mechanisms and life cycle operations Developing modern service architectures around streaming data and event pipelines. Working with infrastructure domain experts on true, zero touch deployment solutions and utilizing best of breed high performance computing management solutions. Be a proactive problem solver, looking out for new opportunities to improve our services and customer experience. Communicate readily with your peers across the organization, b
Jobs in United States
Infrastructure Engineer in United States
1,475 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
Job Requisition ID # 26WD100148 Technical Account Manager – Civil Infrastructure Position Overview Autodesk Customer Success is looking for a highly motivated Technical Account Manager to join our Proactive Support organization. In this role, you will support strategic enterprise customers across the Civil Infrastructure sector. The Technical Account Management team partners with strategic enterprise customers to maximize the value of their Autodesk investment by providing proactive technical guidance, helping customers maintain technically healthy production environments, and enabling them to achieve their business objectives. As a trusted technical advisor, you will build long-term relationships, proactively identify technical risks, and provide recommendations that improve solution performance, operational health, and long-term success. We are seeking professionals with deep experience working with Civil Infrastructure design and engineering workflows in enterprise environments. This role supports organizations delivering transportation, land development, utilities, water, and other infrastructure projects, helping them maximize the value of their Autodesk investment through proactive technical guidance and strategic partnership. Responsibilities Serve as the trusted technical advisor for a portfolio of strategic enterprise customers. Establish and maintain strong customer relationships through proactive technical engagement and strategic guidance. Partner with Customer Success Managers (CSMs) and Technical Adoption Specialists (TASs) to execute
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We are seeking a mid-level Infrastructure Vulnerability Management Engineer with a strong background in Cloud Security, DevSecOps, and Infrastructure-as-Code (IaC). In this role, you will bridge the gap between security, compliance, DevOps, and Platform engineering teams. You will identify infrastructure misconfigurations, secure multi-cloud environments, and manage continuous vulnerability lifecycles across cloud workloads, containers, and data repositories to satisfy strict regulatory compliance frameworks. You will also serve as a technical infrastructure responder during security incidents, deploying real-time cloud or network countermeasures to protect our production ecosystem. What You'll Do Core Responsibilities Infrastructure Scanning & Triage: Perform continuous security scanning across our cloud posture and workloads. Review, validate, and prioritize flaws and misconfigurations based on CVSS scores, real-world exploitability, and infrastructure network exposure. Posture Management & Visibility : Own and optimize Cloud Security Posture Management (CSPM), Kubernetes Security Posture Management (KSPM), and Data Security Posture Management (DSPM) tools to ensure uniform compliance, prevent data leakage, and maintain hardened baselines. Infrastructure-as-Code (IaC) Security: Configure, tune, and embed automated IaC security scanning tools into CI/CD pipelines to identify architectural risks (e.g., overly permissive IAM, public S3 buckets/Cloud Storage) before they are deployed to production. Workload & Container Security: Manage the continuous vulnerability scanning lifecycle for container images, registries, and Virtual Machines (VMs), partnering with SRE and Platform teams to build aut
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Core Platform team provides the foundational infrastructure that powers all engineering at Vanta. We're expanding upmarket to support enterprise customers, which requires strategic investment in platform systems that ensure security, reliability, and developer productivity at scale. As we expand upmarket to support enterprise and regulated customers, we’re investing heavily in platform capabilities that scale securely while reducing cognitive load for product teams. As the Engineering Manager, Core Platform at Vanta, you'll own the foundational infrastructure that every engineer builds on, ensuring it scales with company growth while remaining fast, simple, and reliable. This team’s ownership spans shared services infrastructure, observability and monitoring, datastore management, and async work systems. Our Engineering Managers develop and grow high-performing teams that deliver significant value to our customers and enable our business to scale. This role sits at the intersection of technical architecture and team development, with real authority to set direction and grow a world-class platform team. Visit our Vanta Engineering Blog to learn more about what our team is working on! What you’ll do as an Engineering Manager at Vanta: Lead and grow high-performing platform engineerin
From $196.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Sr. Studio Software Engineer for Roblox Studio Platform, you will be a key contributor to the evolution of Roblox Studio, the primary IDE for making massive multiplayer online games on the Roblox platform. Studio provides the mission-critical tools for 3D modeling, animation, and the complete SDLC for millions of developers. We are looking for engineers who thrive on an adventure into the unknown and have experience across various systems, Operating Systems, Game Engines, and Application Frameworks. You’ll tackle projects involving: Core User Features: Architecting application frameworks, windowing systems, and code generation. “AI Native” Features: Pioneering scalable systems that extend to complex, agentic use cases. Foundational Architecture: Driving OS integration, extensibility, and customizability at the deepest levels. Central Backend Systems: Engineering the infrastructure to power a consistent, high-performance UX. Design Evolution: Collaborating with UX designers to translate Studio into a modern, consistent design language. You Will: Design and execute the technical direction to drive the future extensibility and adaptability of the application. Own and deliver complex techn
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Team Overview The TechOps team is the technical foundation that is a major stakeholder in keeping the company running. We own the systems, tools, and infrastructure that every Plaid employee depends on, from identity and endpoint management to help desk support, office infrastructure, and the internal tooling that powers day-to-day productivity. What sets us apart from a traditional IT team is how we approach our higher-level goals. We treat corporate infrastructure like an engineering problem: configuration lives in code when possible, endpoint provisioning is automated, access assignment is self-service, and we're always looking for ways to make our systems more reliable and our support burden smaller. We're a small team with broad ownership and high standards. We work closely with Security, Engineering, and People teams to make sure Plaid's internal environment is secure, scalable, and ready for where the company is going, whether that's a new office, a new compliance requirement, or a new way of working enabled by AI tooling. Role Overview In this role, you'll take part in our on-call rotation for a few hours each week, but this isn't just a break/fix IT position. It's a chance to build, improve
From $128K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. Join Our Team… GoDaddy's Global Storage Engineering team operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you're passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you'll work on infrastructure that few engineers ever experience. At GoDaddy, Ceph isn't a side project — it's a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering. As a Senior Site Reliability Engineer, you'll be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You'll tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide. What You'll Get to Do… Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage wor
$155K – $207K/yr
About Taskrabbit: Taskrabbit is a marketplace platform that conveniently connects people with Taskers to handle everyday home to-do’s, such as furniture assembly, handyman work, moving help, and much more. At Taskrabbit, we want to transform lives one task at a time. As a company we celebrate innovation, inclusion and hard work. Our culture is collaborative, pragmatic, and fast-paced. We’re looking for talented, entrepreneurially minded and data-driven people who also have a passion for helping people do what they love. Together with IKEA, we’re creating more opportunities for people to earn a consistent, meaningful income on their own terms by building lasting relationships with clients in communities around the world. Taskrabbit is a hybrid company with employees distributed across the US and EU and a Built In — Best Places to Work (2022, 2023, 2024, 2025) continually ranked across multiple national and regional categories. Join us at Taskrabbit, where your work will be meaningful, your ideas valued, and your potential unleashed! This role operates on a hybrid schedule requiring two days of in-office collaboration per week. About the Role Reporting to the Infrastructure Director of Engineering, you are a pragmatic security leader who acts as a business enabler rather than a gatekeeper. You thrive in a high-velocity, non-regulated environment where you must define "what good looks like" from scratch and drive security progress with a blend of strong planning skills and strong cross-functional partnerships. You are a technical player-coach who manages the "how" behind engineering initiatives, providing direct guidance to a lean team which needs to navigate trade-offs between quality, depth, and delivery volume. You are just as comfortable planning core initiatives as you are writing a proposal for security program improvements. You have a proven track record of leading a team to achieve compliance with a security framework from scratch (e.g., CIS, SOC2, P
About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning. The team’s work spans four areas. First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early. Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy. Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains. Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with. About the Role This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory. This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. In this role, you will: Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems. Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure. Improve the reliability and efficiency of RL trai
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. Team Overview Plaid's TechOps team is the technical foundation that is a major stakeholder in keeping the company running. We own the systems, tools, and infrastructure that every Plaid employee depends on from identity and endpoint management to help desk support, office infrastructure, and the internal tooling that powers day-to-day productivity. What sets us apart from a traditional IT team is how we approach the higher level goals we have to improve our systems and tooling over time. We treat corporate infrastructure like an engineering problem: configuration lives in code when possible, endpoint provisioning is automated, access assignment is self-service, and we're always looking for ways to make our systems more reliable and our support burden smaller. We've built real momentum in that direction, and we're investing in the people and tools to take it further. We're a small team with broad ownership and high standards. We work closely with Security, Engineering, and People teams to make sure Plaid's internal environment is secure, scalable, and ready for where the company is going whether that's a new office, a new compliance requirement, or a new way of working enabled by A
$151K – $201K/yr
At Playlist, life's richest moments happen when people step away from screens to move, connect, explore, and play. We're building the definitive platform for intentional living, connecting people with inspiring experiences in fitness, wellness, and beyond. With popular brands like Mindbody and ClassPass, Playlist empowers businesses and individuals, making it effortless for aspirations to become actions. Join us in reshaping technology's role to foster meaningful, real-world connections. Mindbody equips wellness entrepreneurs with technology to support thriving businesses and create exceptional experiences. Innovation and curiosity drive our culture, connecting businesses and individuals through cutting-edge solutions. Join us if you're passionate about enhancing wellness through technology. You’ll likely spend time working on: Empowering your engineering team to deliver high levels of technical quality and business impact. Recruiting a diverse team of engineers recognized as some of the best in the industry, from outreach to closing. Managing a team of engineers and facilitating their career development through challenging responsibilities and constructive feedback. Actively participating in strategic direction and product decisions for your team. Contributing to engineering-wide initiatives as a member of the Playlist engineering management team. About the right team member: You are an experienced engineer who has successfully transitioned into a people leader, passionate about building and operating reliable, scalable cloud platforms and developer infrastructure that empower engineering teams. You possess extensive experience in at least one aspect of developing a world-class cloud platform and can effectively manage a high-performing team throughout the entire process. You’ll thrive in this role with experience in: 3+ years of direct management experience with an engineering team of at least 6 members with skills acro
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Operations Program Manager (OPM) to serve as the single-threaded operational leader for new hardware introductions (NPI) and production ramps across OpenAI’s AI infrastructure systems. This role combines hands-on execution with strategic ownership. You will be responsible for defining the operating model, aligning cross-functional stakeholders, setting the critical path, making informed tradeoffs, escalating decisively, and ensuring hardware programs deliver on schedule, quality, cost, and scalability. Success in this role requires comfort operating in ambiguity, influencing without authority, and driving alignment across internal teams and external partners—while keeping eyes firmly on long-term system scalability and repeatability. In this role, you will: Strategic & Leadership Ownership Act as the single-threaded owner for operational readiness across NPI and ramp, accountable for outcomes from early bring-up through sustained production Translate OpenAI’s infrastructure strategy and engineering objectives into clear operating plans, execution priorities, and decision frameworks Drive alignment across Engineering, Operations, Strategic Sourcing, Finance, Capacity Planning, and Executive stakeholders by framing tradeoffs, risks, and recommendations Proactively identify inflection points where decisions or investments are required to protect long-term scale, reliability, or cost targets Influence operational strategy with manufacturing par
From $204K/yr
The opportunity Datadog’s Infrastructure products help engineers understand and operate the systems their applications depend on. Our customers work in complex environments like Kubernetes and serverless, where infrastructure changes constantly, information is dense, and decisions about reliability, performance, and cost are closely connected. We’re looking for a Staff Product Designer to join Modern Compute, with an initial focus on Containers Autoscaling. Autoscaling helps engineering teams make better decisions about how their applications and infrastructure use resources. Designing these experiences requires making deeply technical systems understandable, helping customers act with confidence, and fitting into the tools and workflows they already use. The team is rethinking how workload and cluster autoscaling come together as a more coherent product experience. This includes how customers get started, understand recommendations, evaluate value, and safely apply changes across their environments. The work also connects to other parts of Datadog, including observability, Cloud Cost Management, permissions, and AI-assisted workflows. As a Staff Product Designer, you will help define that direction and lead the work from early problem framing through shipped product. You will partner closely with product and engineering, bring a high level of interaction and visual craft to complex workflows, and help raise the quality of design across Modern Compute. At Datadog, we place value in our office culture, the relationships and collaboration it builds, and the creativity it brings to the table. We operate as a hybrid workplace to help our Datadogs find a work-life rhythm that works for them. What you’ll do Lead end-to-end product design for Modern Compute, initially focused on our Autoscaling product. Help define the product direction for an area that is still evolving, from early framing and exploration through detailed design and delivery. Design clear, trustwort
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. You will build the layer that makes Vanta's data actually useful: a single interpretive layer that reads across every source of truth and makes them queryable, reasoned-over, and genuinely illuminating — for EPD, for GTM, and for anyone in the org trying to understand what's happening and why. The EPD Systems team is building the infrastructure Vanta's engineering, product, and design organization depends on to understand itself. We're constructing three interlocking layers: sources of truth at the foundation, an operating system that makes delivery legible, and the intelligence layer that reasons across all of it. This role owns the intelligence layer — the one that doesn't exist yet. This is a builder role. You will ship working things yourself — prototypes, internal tools, agent workflows. You will not hand specs to someone else and wait. Communication isn't a separate deliverable; it's how you learn what to build. What you’ll do as a Senior Product Builder at Vanta: Build the intelligence layer: a cross-source interpretive layer that reads across Vanta's sources of truth and makes them queryable and reasoned-over by EPD leadership, GTM, and beyond Define what to build: scope the problem yourself, make explicit tradeoffs about what to defer, and own the sequence of what gets built and when Ship working things yourself: prototypes, internal tools, agent workflows — using AI as part of how you work, not what you report on Understand the organization: go to the teams whose decisions this layer will serve — GTM, G&A, EPD — and come back knowing what they can't answer today Catch and address AI-specific quality problems — rel
About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: Modal builds the infrastructure that lets engineers run AI workloads without the usual pain. To do this well, we need exceptional people – and that’s where you come in. As the first dedicated GTM recruiter on our Talent team, you’ll own sales, GTM, and other G&A searches end-to-end. You’ll work closely with our Head of Talent, founders, and GTM leads to shape how we hire and help bring in the people who will define what Modal becomes. What you’ll do: Drive full-cycle recruiting for key hires across GTM and G&A functions (sourcing, pitching, guiding interviews, and closing candidates) Partner with GTM leaders to understand the real work and calibrate on what great looks like Help set our hiring bar and how we evaluate talent Execute creative top-of-funnel strategies that resonate with a strong community of experienced GTM talent Deliver a fast, respectful, h
Other cities to consider
More places hiring for this role
Get new infrastructure engineer jobs in United States by email
Daily job updates · Unsubscribe anytime