Jobs in United States

Infrastructure Security Engineer in United States

1,531 active opportunities · Updated October 2026

Explore current infrastructure security engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of a best-in-class family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from a diverse group of backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a senior validation lead engineer to lead at-scale rack validation efforts for next-generation AI hyperscale systems. This role focuses on post-silicon system validation across the full lifecycle, ensuring functional, electrical, and thermal performance meets product objectives. You will own end-to-end blade and rack validation including planning, development, execution, and debug while collaborating across firmware, systems, and hardware teams. The Team The Rack Validation team is responsible for ensuring system readiness and quality at scale. The team works cross-functionally with firmware, silicon, and system engineering teams to validate complex AI compute platforms. Responsibilities and Duties Lead post-silicon validation of AI compute blades and racks including test planning, development, and automation. Drive provisioning and integration of system components (SoC FW, BMC, RMC, OS) for rack-level readiness. Own execution against program achievements and report validation progress and risks. Triage test failures, collect debug data, and collaborate on root cause analysis. Track

PythonCI/CDLinuxAI
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonAIExcelHR
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Staff -Power and Performance Validation Engineer About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to senior leadership within Architecture and Validation, the Power and Performance Validation Lead will drive validation strategy and execution for advanced AI compute silicon and systems. The role is responsible for leading power, thermal and performance validation activities across pre-silicon and post-silicon environments to ensure products meet efficiency, reliability and scalability expectations. This role requires strong technical expertise and collaboration across multiple engineering disciplines to deliver robust validation methodologies, scalable automation frameworks and actionable performance insights. The Team The Power and Performance Validation team sits within the Architecture and Validation organisation and is responsible for validating the performance, efficiency and thermal behaviour of Graphcore silicon and systems. The team supports the full product lifecycle, from early architectural modelling through to first silicon bring-up, characterization and production readiness. Engineers work closely with cross-functional teams globally to debu

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced Silicon Test Engineer to join our Product Test and Diagnosis Department (PTD). This is a pivotal role and will involve building a team of engineers to develop System Level Test (SLT) capability within the company. Working closely with a cross-functional team you will implement SLT tests for a family of next generation AI Processors. The ideal candidate should have a focus on quality and demonstrate a good understanding of the importance of production test on the success of a product. T hey will have a proven Functional Test or ATE Test Engineering background, and will have a pragmatic, hands-on and flexible approach to a fast-changing environment. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Managing a team of

PythonGitAIGo
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary Reporting to the Quality leadership within Manufacturing Operations, the Senior Reliability Scientist is responsible for leading reliability activities across complex, high-performance systems. Working closely with established reliability experts and cross-functional teams, this role uses experimental data and advanced modelling to inform design decisions, validate product reliability and optimise serviceability strategies, including spares provisioning. The Team The Quality team within Manufacturing Operations is responsible for ensuring product robustness, reliability and lifecycle performance across Graphcore’s hardware portfolio. The team includes experienced reliability specialists and works closely with technology research, chip, board, system design, platform and operations teams to translate reliability insights into actionable improvements across the product lifecycle. Responsibilities and Duties: · Define and refine reliability requirements across silicon, board and system levels, working in partnership with research and design teams · Apply ad

AIGoExcelSEM
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will contribute to the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with hardware, software, validation, and manufacturing teams, you will develop test content, automation, diagnostics, and characterization capabilities that support silicon bring-up, yield learning, and manufacturing deployment. The ideal candidate will have strong technical foundations in semiconductor test and validation, excellent debug skills, and a passion for improving product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Develop and maintain SLT test content, automation, d

PythonGitAIExcel
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems. The Team The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and infe

PythonAIGoDevOps
P
📍 New York, New York, United States· Full-time· Remote
✓ High-confidence listingCompany trend -73.5%
Quick readStrong listing-quality and freshness signals

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Data team within Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection products, leveraging Plaid’s unique network data to identify and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production serving and monitoring—building reliable, scalable systems that deliver high-quality fraud detection as we grow to support hundreds of customers. As a Senior Machine Learning Engineer, you will own the development of high-performance feature computation and online inference pipelines that power production machine learning systems at scale. You’ll build robust observability, monitoring, and automated debugging capabilities, while leveraging AI-assisted tools to investigate complex system behavior and maintain high reliability. You’ll partner closely with ML Infrastructure, Data Science, and Product teams to execute critical technical initiatives and deliver scalable, high-impact ML solutions. Responsibilities: Build and scale machine learning systems that power a rapidly growing fraud detection product in a fast-paced environment. Solve complex technical challenges at the intersect

PythonAWSMachine LearningAI
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $156K/yr

Quick readStrong listing-quality and freshness signals

As a Product Manager – IaC Detection, you will define, build, and launch capabilities that proactively detect infrastructure issues in code (e.g. Terraform, Helm) before they can be deployed into production and escalate into production incidents. The Infrastructure Monitoring team has pioneered shift-left detection in the industry with Bits Infrastructure Operations , and we’re looking for a Product Manager to expand this capability to a broader set of use cases Customers (and thus developers) are increasingly standardizing on IaC tools to deploy and maintain ever-growing infrastructure in the cloud. At the same time, SREs and Infra teams struggle with an increasing number of production incidents. By shifting-left and identifying high-impact infra changes before they are deployed, we help reduce production incidents, reduce waste, and free up SRE time to focus on value-added tasks. You will own the roadmap to expand IaC detection to a broader set of use cases, including cost detection, blast radius impact, as well as configuration changes on infrastructure powering applications like nginx, postgres and more. You’ll partner closely with Engineering, Design, and customers to build and iterate on the roadmap, build product market fit, drive customer adoption (including internal usage), and focus on coverage and correctness of the AI system. This is an opportunity to lead an initiative at the intersection of AI, infrastructure operations, and autonomous observability. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Lead the product roadmap for IaC Detection, enabling customers to proactively detect and catch high-impact infrastructure and configuration changes before they are deployed into production and escalate into incidents. Define the end-to

GitAIRustTerraform
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -83%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE Baseten's Compute org is in hyper growth. As it scales, the systems and workflows that keep supply and demand balanced across our GPU fleet need to get more sophisticated, and this role exists to make sure they do. Compute sits at the center of how Baseten allocates, forecasts, and manages the capacity that powers every customer inference request. The team that supports this work, C3, runs on a mix of internal tooling, manual processes, and systems that haven't fully kept pace with the scale of the problem. This role exists to close that gap. You'll design, build, and ship AI-powered workflows that give the Compute and C3 teams real leverage, automating the manual, repetitive, and error-prone parts of the capacity lifecycle so the team can focus on judgment calls that actually need a human. We want someone who can walk in, audit what exists today, identify what's missing or broken, and start shipping fast. You know when to reach for an existing internal tool and when to build something custom in Claude Code. You think two to three steps ahead about how the thing you build today fits into the broader capacity systems architecture tomorrow. And you bring a point of view on our stack, on what we should be building, and on where AI can do something existing tooling simply can't. RESPONSIBILITIES Ship AI-powered workflows for Compute and C3 : build the agents and automations that give capacity analysts, ops leads, an

Machine LearningAIGoRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -84.1%
Quick readStrong listing-quality and freshness signals

About the Team The Emerging Products team is a lean, high-output product lab group that builds products at the forefront of model capabilities. We collaborate across all teams within the company, from research and infrastructure to consumer products. The team is responsible for identifying new product opportunities, building them quickly, dogfooding them internally, and then launching the successful products to users. We use data, user research, and analytics to inform our ideas, and make decisions on what experiments are worth iterating, stopping, or scaling. About the Role We’re looking for a senior, product-minded software engineer to own ambiguous 0-to-1 work from idea through prototype, validation, and handoff. This is a full-stack role with a strong frontend and product emphasis: you will build the interfaces and supporting backend systems needed to test new experiences quickly, while making sound architectural choices that enable successful concepts to scale. This role is based in our Mission Bay office in San Francisco. In this role, you will: Build and ship high-quality, product experiments across the full stack. Turn ambiguous user needs and emerging technical capabilities into testable product concepts, using research and metrics to guide iteration. Own technical direction for 0-to-1 projects, balancing speed, reliability, and a clear path from prototype to scalable product. Partner closely with design, product, research, and engineering teams to dogfood, evaluate, launch, and transition successful experiments. You might thrive in this role if you: Have a track record of building and shipping end-to-end products in fast-moving, startup, founder-led, growth, or other high-ownership environments. Bring strong frontend engineering skills and enough backend and systems depth to make sound full-stack architectural decisions. Pair product intuition with evidence, using user research and product data to identify opportunities and make pragmatic tradeoffs. Operat

AWSRestAIRust
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -83%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE We’re looking for a high-performing strategic finance professional to join our growing GTM Finance team. Our business grows with our customers' usage, which makes the finance function highly strategic at Baseten: growth, pricing, margin, and capacity decisions are business model decisions. You'll sit at the center of them, partnering directly with GTM leadership and reporting into a finance team with a seat at the table for the calls that shape the company's trajectory. This role is ideal for someone with 3 to 7 years of experience across strategic finance, investing, and/or investment banking who wants broad exposure to company-building inside a fast-scaling AI infrastructure company. Experience at a usage-based software company is a plus. RESPONSIBILITIES Own financial planning, forecasting, and budgeting processes for the GTM org Build and maintain financial models across revenue, S&M spend, headcount, and strategic bets Analyze the metrics that define a usage-based business – ARR, gross margin, consumption trends, retention, and GTM efficiency Partner with GTM leaders to set targets, evaluate growth initiatives, shape pricing, and design sales compensation Help prepare board materials, investor updates, and fundraising analyses Improve financial reporting, dashboards, and operational rigor so our infrastructure scales as fast as our revenue Work cross-functionally to turn ambiguous business questions int

SQLRestMachine LearningAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -84.1%

About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. The Power Execution team owns the strategy and execution required to secure reliable, scalable, and economically resilient power for OpenAI’s global data center portfolio. The team sits at the intersection of commercial, technical, policy, legal, and operational work, partnering across OpenAI and with utilities, grid operators, regulators, counterparties, and public-sector stakeholders. About the Role The Energy Regulatory Lead will own energy regulatory strategy and execution for OpenAI’s infrastructure growth. This role will be the primary bridge between the Power Execution team and Public Policy and Government Affairs on energy regulatory matters, ensuring that OpenAI’s external engagement is grounded in project realities and that changing policy and regulatory conditions are translated into actionable infrastructure decisions. This is an individual contributor lead role and does not have direct reports initially. The role combines portfolio-level regulatory positioning with transactional regulatory work: evaluating jurisdictional pathways, supporting utility and energy transactions, coordinating approvals and filings, and helping project teams navigate tariffs, interconnection, load-service requirements, market rules, and regulatory risk from diligence through execution. In this role, you will: Develop and maintain OpenAI’s energy regulatory strategy across priority U.S. markets and, as needed, emerging geographies for infrastructure expansion. Coordinate closely with Public Policy and Government Affairs to shape energy regulatory priorities, engagement plans, messaging, and positions before utilities, public utility commissions, grid operators, state energy offices, and other relevant policymakers. Translate project requirements—load size, timing, reliability, cost, carbon, and expansion needs—into clear regulatory objectiv

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend 0%

From $204K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. About the role: The Customer Success Executive - Enterprise role will play a critical part in the exciting growth and retention trajectory of the Enterprise Sector at Okta. You will report to the Regional Manager, Customer Success and be responsible for complementing Okta’s innovations, best practices, and capabilities with our valued customers’ business objectives and priorities. As a strategic and trusted advisor, you will shift customer mindsets, create demand for our solutions, and drive higher business value and executive alignment between Okta and our customers. The ideal candidate embodies a high-performance mindset—consistently raising the bar, turning action into traction, thriving in change, and moving fast to simplify and repeat. It requires complementing Okta’s innovations, efficiencies and capabilities with our valued customers’ business objectives and priorities thereby driving higher business value and executive alignment between Okta and our customers, but it also demands a leader who can navigate these constraints with urgency and an obsession for concrete outcomes. Strong problem-solving, orchestration, and consultative skills are necessary for navigating challenges, finding innovative solutions, and winning as a team. In this role, you will: Deliver Customer Value: Develop and nurture strong customer and C-Level relationships to understand their business goals and needs, ensuring retention, happiness, and a significant return on inv

AWSRestMachine LearningAI
S
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -80.6%

About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About this role At Sentry, Finance plays a critical role in helping the company understand where we're investing, what we're getting in return, and where we can operate more effectively. We are looking for an FP&A Manager who will help build Sentry's FP&A function while focusing on establishing best-in-class FinOps practices. You'll partner across the business on forecasting and financial analysis, with a significant focus on cloud infrastructure, AI costs, vendor spend, and other areas where better financial visibility can translate directly into better decisions. You will be expected to understand the economics behind our infrastructure, build the financial models and reporting needed to manage it, identify meaningful opportunities, and work with Engineering and other partners to turn those insights into action. You'll also operate as a generalist FP&A partner and play an important role in building the processes, models, and operating rhythms that Finance will use as Sentry grows. You will report to the Director of FP&A. In this role you will Own forecasting, reporting, and financial analysis for significant areas of Sentry's operations, with particular emphasis on cloud infrastructure, AI-related spend, software, vendors, and other major cost categories. Partner closely with Engineering and infrastructure leaders to understand the drivers of cloud and AI spend, translate technical consumption into financial forecasts, and identify opportunities to improve unit economics and gross margin. Develop reporting that makes cloud and infrastructure costs understandable and actionable, including trends, cost alloca

RestAIGoRust
🔔

Get new infrastructure security engineer jobs in United States by email

Daily job updates · Unsubscribe anytime