We are investing in agentic AI and need a Senior AI Engineer to lead the design and delivery of these systems. This is a foundational hire: you will own both the agent-facing workstreams — pipelines, orchestration, conversational interfaces — and the underlying context layer that makes them reliable, including memory management, knowledge graph integration, and retrieval infrastructure. You will work closely with data engineers, project leads, and client stakeholders, and play a key role in shaping how Lynx builds and ships AI solutions at scale. What This Involves: Lead the architecture and delivery of agentic AI systems end-to-end: agents, orchestration, tool use, and multi-step reasoning workflows. Own the context layer: design and implement memory architectures (episodic, semantic, working memory) and integrate GraphRAG and knowledge graph retrieval into agentic pipelines. Build robust RAG systems — including vector retrieval, graph traversal, and hybrid search — and ensure retrieval quality through evaluation frameworks. Translate client requirements into technical designs, presenting approaches and trade-offs to both technical and non-technical stakeholders. Define standards and reusable patterns for agentic AI development that other engineers at Lynx can build on. Set up observability, evaluation, and monitoring pipelines to ensure AI systems perform correctly in production. Requirements: 5–8 years of software or ML engineering experience, with at least 2–3 years building LLM-based or agentic AI systems in production. Deep hands-on experience with agentic frameworks (LangChain, LlamaIndex, AutoGen, CrewAI, or similar) and LLM APIs (OpenAI, Anthropic, etc.). Strong understanding of agent design patterns: ReAct, planning loops, tool use, multi-agent coordination, and memory architectures. Practical experience with GraphRAG or knowledge graph-based retrieval (e.g., Neo4j, Microsoft GraphRAG) and vector databases (Pinecone, Weaviate, Qdrant, etc.). Proficiency in
Jobs in United States
Lead Infrastructure Software Engineer in United States
2,434 active opportunities · Updated October 2026
Showing
15 jobs
Explore current lead infrastructure software engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About Graphcore At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale.As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence. Job Summary We are looking for an experienced System Level Test Engineer to join our Product Test and Diagnosis Department (PTD). In this role, you will lead the development and deployment of System Level Test (SLT) solutions for next-generation AI processors. Working closely with cross-functional teams, you will contribute to the design and implementation of SLT hardware, software, automation, and characterization solutions that support silicon bring-up, yield learning, manufacturing readiness, and production deployment. The ideal candidate will possess strong technical depth in semiconductor test and validation, a passion for solving complex engineering challenges, and a strong focus on product quality and manufacturability. The Team The Product Test and Diagnostics team’s role is to detect and manage hardware defects that arise from the manufacture and use of our products. This covers chips, boards and finished systems and takes place both in the manufacturing sites and in the field. Responsibilities and Duties Lead development and deployment of SLT hardware and software solutions supporting silicon bring-up, charac
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in the company. They build the internal control platforms, configuration systems, traffic routing, and automation that let product engineers operate services safely at scale without needing deep infrastructure expertise themselves. As the Engineering Manager for Infrastructure Engineering, you'll lead a team of engineers building the tools that power Sentry's growth: internal admin and change management tools, configuration automation, and the routing layer that underlies Sentry's architecture. You'll be responsible for technical vision, team health, system reliability, and partnership with engineering teams across the company who depend on your team's tools every day. You'll work closely with leaders across Infrastructure, Platform, and Production Engineering to shape how Sentry scales its operational model as the company grows. In this role you will Lead a team of engineers building the internal control platforms that every engineering team at Sentry relies on to operate services safely. Drive the evolution of Infrastructure Engineering's platform, including configuration management, traffic routing and environment controls Own the team's technical direction, contributing to key decisions on API architecture, internal tooling design, and automation frameworks. Nurture and grow engineers at different levels, providing support through coaching, mentorship, and career development. Foster an inclusive, high-performing team culture focused on ownership, learning, and delivery. Partne
$220K – $450K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the role Sentry's Infrastructure Engineering team is what makes operating Sentry simple, safe, and seamless for every other engineering team in the company. They build the internal control platforms, configuration systems, traffic routing, and automation that let product engineers operate services safely at scale without needing deep infrastructure expertise themselves. As the Engineering Manager for Infrastructure Engineering, you'll lead a team of engineers building the tools that power Sentry's growth: internal admin and change management tools, configuration automation, and the routing layer that underlies Sentry's architecture. You'll be responsible for technical vision, team health, system reliability, and partnership with engineering teams across the company who depend on your team's tools every day. You'll work closely with leaders across Infrastructure, Platform, and Production Engineering to shape how Sentry scales its operational model as the company grows. In this role you will Lead a team of engineers building the internal control platforms that every engineering team at Sentry relies on to operate services safely. Drive the evolution of Infrastructure Engineering's platform, including configuration management, traffic routing and environment controls Own the team's technical direction, contributing to key decisions on API architecture, internal tooling design, and automation frameworks. Nurture and grow engineers at different levels, providing support through coaching, mentorship, and career development. Foster an inclusive, high-performing team culture focused on ownership, learning, and delivery. Partne
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. An overview of this role As a member of the Infrastructure Security Team within the Product Security Department , you will work with teams across GitLab to ensure that the components that comprise our public cloud infrastructure are built from the beginning with resiliency and set security expectations that our customers rely on to power their DevSecOps goals. As a Staff Security Engineer, you will serve as a technical lead across the topics the Infrastructure Security team owns, including our SaaS Platforms (e.g. GitLab Dedicated, Cells) and Self-Managed offerings. You will define the technical direction for how the team approac
The Engineering Lead Analyst – SonarQube & Code Quality Engineering is a senior-level engineering role responsible for leading static code analysis, automated code quality governance, security vulnerability remediation, and AI-augmented developer enablement across enterprise software delivery pipelines. In this role, you will champion software reliability, maintainability, clean-coding standards, and automated quality gates. You will partner with development teams, system architects, and platform engineering to integrate and manage enterprise-scale code quality platforms (such as SonarQube) both on-premises and in cloud/SaaS environments. Additionally, you will drive modern engineering practices by embedding Behavior-Driven Development (BDD) within your own software delivery and leveraging Agentic AI workers and Model Context Protocol (MCP) architectures to optimize developer experience, streamline code governance, and boost engineering velocity. Key Responsibilities 1. Code Quality & Static Analysis Platform Ownership Lead the architecture, deployment, administration, and continuous enhancement of enterprise Static Application Security Testing (SAST) and Code Quality platforms (e.g., SonarQube , DeepSource, Codacy, Semgrep). Configure, calibrate, and enforce automated Quality Gates, code rulesets, technical debt calculation models, and code-coverage baselines across multi-language enterprise repositories. Oversee version upgrades, patching, high availability, and operational maintenance for on-premises and SaaS/cloud-hosted code quality infrastructure. 2. CI/CD & Pipeline Integration <li style=
Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) **Sign on Bonus Potential** Company: The Boeing Company The Boeing Company’s Specialized United States Infrastructure Operations organization is currently seeking a Cloud Infrastructure Administrator (Mid-Level, Senior or Lead) to join the team in Berkeley, MO; Seattle, WA; or Daytona Beach, FL . The Infrastructure team is seeking an experienced cloud infrastructure professional to help design, build, and sustain the foundational cloud environment supporting critical program needs. In this role, the selected candidate will help establish and operate secure, scalable, and resilient cloud infrastructure environments in Microsoft Azure to enable enterprise applications, software toolchains, and digital engineering workloads. As both an individual contributor and technical leader, this position will work across network, computer, storage, identity, security, and automation domains to deliver repeatable cloud infrastructure patterns and operational excellence. This role is focused on infrastructure operations, sustainment, automation, and reliability, rather than application software development. Position Responsibilities: Design, implement, and maintain Microsoft Azure-based infrastructure solutions including networking, compute, storage, identity integration, and supporting services Develop and maintain Infrastructure as Code (IaC) and configuration automation solutions using Terraform, Ansible, PowerShell, and Bash Implement cloud policies to enforce security, ensure regulatory compliance, and manage user access Build repeatable landing zones and cloud infrastructure patterns that support mul
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We're looking for an AI-native product growth lead to own Replit's entire paid acquisition and lifecycle marketing engine as the single DRI. This is an IC role for someone who builds systems, not just campaigns. Someone who uses AI and automation aggressively to do what traditionally requires an entire team. You don't need an engineering background. You need the instinct to build: when you see a repetitive task, you reach for Replit or an API before you reach for a spreadsheet. You'll have data science and data engineering partners for infrastructure and modeling, and brand marketing will set creative direction. Your job is to turn that direction into a closed-loop optimization system: a self-improving system that turns performance data into better creative, optimizes spend across channels in near-real-time, and compounds every insight so nothing learned is ever lost. Each cycle, the system gets smarter. That's the engine you'll build and own. The ideal candidate has deep channel expertise across paid search, paid social, and lifecycle, but their real edge is building AI-powered workflows that scale creative production, automate measurement, and compound institutional knowledge. You Will Own full-funnel performance marketing across paid search & social, ASO, and lifecycle Build a self-improving creative engine: AI-driven ad generation, testing, and iteration that scales without scaling headcount Maintain a persistent knowledge layer so every experiment, result, and creative insight compounds automatically into the next cycle Own conversion signal quality end-to-end: right events, right audiences, right attribution, partnering with DE on the infrastructure Run a structured experimentation program wher
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. We are looking for a Security Operations Lead (SOC Lead) to build, mature, and operate our 24/7 detection and response capabilities across a modern cloud-native and AI-driven environment. This role leads the global SOC function—monitoring, SIEM ownership, detection engineering, alert triage, and operational readiness—while also evaluating and integrating emerging AI-based SOC products and autonomous response platforms . You will oversee monitoring across multi-cloud environments (GCP primary, AWS/Azure secondary), Kubernetes, SaaS services, endpoints, developer tools, and AI workloads . You’ll collaborate closely with Cloud Security, Compliance/GRC, SRE, Platform Engineering, IT/Endpoint teams, and AI Infrastructure to ensure our detection strategy scales and stays ahead of evolving threats. This is a hands-on leadership role perfect for someone who wants to shape the SOC of the future while solving complex challenges in a high-scale AI setting. What You’ll Do SOC Leadership & 24/7 Monitoring Lead, mentor, and scale a global SOC team responsible for 24/7 monitoring, alert intake, triage, correlation, and escalation. Build operational rigor: processes, runbooks, SLAs, metrics, and quality standards for high-scale environments. Cover monitoring across: Cloud infrastructure (GCP, AWS, Azure) Kubernetes/GKE/EKS/AKS clusters SaaS platforms (Google Workspace, GitHub, Slack, Okta, etc.) Endpoints (macOS, Linux, Windows) including EDR/XDR telemetry Developer platforms + CI/CD pipelines AI/ML systems and model-serving workflows AI-Based SOC Integration & Innovation Evaluate, adopt, and integrate AI-native SOC technologies for triaging, detection, and correlation Identify opportunities to automate triage, investigations,
From $220K/yr
About Sentry Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building. Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future. About the team The Billing team sits at the intersection of product, finance, and infrastructure. They're responsible for ensuring every observable event—errors, logs, traces, tokens—gets accurately measured, priced, and billed. Their work directly impacts company revenue and customer trust, requiring distributed systems expertise, attention to financial accuracy, and deep understanding of product usage patterns. The team works cross-functionally with product, engineering, BizOps, marketing, and sales to build systems that enable new products and pricing models. As an Engineering Manager, you’ll lead a team of engineers owning critical workflows such as checkout and invoicing, while also developing new features to help customers manage their spend growth. In this role, you’ll partner across the organization to ensure our customers redeem everything Sentry has to offer and budget for future expansion. In this role you will Strategic Planning & Roadmap: Define and drive the team's roadmap. Align team goals with organizational objectives and contribute to the overall platform strategy. Technical Guidance & Operational Excellence: Provide technical leadership and guidance on complex distributed systems and design. Ensure the team is proactively identifying areas for improvement. Cross-functional Collaboration: Partner closely with business and technical teams to translate business goals into actionable objectives and scalable solutions. Team Leadership & Development: Lead, mentor, and grow a team of talented engineers, including Staff-level engineers. Build a culture of technical excellence, collaboration, continuous
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Role We're hiring a hands-on Engineering Manager to build and lead Replit's Anti-Abuse team from the ground up. This is a foundational 0-to-1 role: you'll define the anti-abuse roadmap, hire a small team of engineers and data analysts, and ship the systems that protect Replit's platform, users, and economics from adversarial actors. You'll partner across Support, Legal, Security, Infrastructure, and the Money and Growth teams to make abuse economically unviable while keeping friction low for legitimate users. Replit sits at the frontier of AI-native abuse. Our platform is a target for phishing and scam hosting, cryptomining, LLM token farming, card and coupon fraud, and increasingly, abuse driven by AI agents themselves. The team you build will define how Replit defends against all of it. What You'll Do Build the anti-abuse roadmap from scratch : Define the threat model, prioritize across abuse vectors (phishing/scam hosting, cryptomining, token farming, payment fraud, AI agent exploitation), and translate it into a shipping plan with clear sequencing and tradeoffs. Design progressive verification and identity infrastructure : Build the "ladder of trust" that gates increasing platform capabilities (referrals, additional credits, access to powerful agent features, Missions) behind escalating verification. This includes a humanity/identity layer that's distinct from user accounts, integrations with KYC-grade verification providers, and the policy engine that decides what level of trust unlocks what behavior. This infrastructure is core not just to promo integrity but to how Replit safely expands agent capabilities over time. Ship as a hands-on EM : Stay in the code. Use the latest AI coding tools (including Rep
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Support Systems Lead at Replit, you'll build and own the systems that let support scale as fast as the product does. You'll own the tools our team works in every day, the technical setup behind self-service and AI-driven support, and the infrastructure that keeps all of it current as Replit ships. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You'll shape how our support systems adapt to new products, new surfaces, and AI-assisted workflows, operating effectively in ambiguity and turning ad hoc fixes into infrastructure the whole team can rely on. You'll combine hands-on technical depth with systems thinking to keep builders moving, whether they get unblocked through self-service, an AI agent, or a person. This is an individual contributor role to start, with room to grow and build out a team as support scales. IN THIS ROLE YOU WILL: Configure and maintain Zendesk and the surrounding support stack, from the day-to-day workflow and automation setup to business rules and permissions, keeping it able to flex and scale as needs change. Build and maintain assignment logic, queues, tagging and taxonomy, and escalation paths, keeping them running cleanly as volume and workflows change. Set up and maintain support tooling, workflows, and access across internal agents, outsourced vendors, and regions, keeping the systems working for each group as the stack changes. Partner with Engineering on the technical requirements for self-service and in-product support surfaces, and build the entry points, help widgets, and routing behind them. Spot the repetitive steps in agent and admin workflows before they become bottlenecks, and build the automations, bulk acti
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. As a Support Systems Lead at Replit, you'll build and own the systems that let support scale as fast as the product does. You'll own the tools our team works in every day, the technical setup behind self-service and AI-driven support, and the infrastructure that keeps all of it current as Replit ships. Replit is at the forefront of AI-driven software development, and how we support customers is constantly evolving. You'll shape how our support systems adapt to new products, new surfaces, and AI-assisted workflows, operating effectively in ambiguity and turning ad hoc fixes into infrastructure the whole team can rely on. You'll combine hands-on technical depth with systems thinking to keep builders moving, whether they get unblocked through self-service, an AI agent, or a person. This is an individual contributor role to start, with room to grow and build out a team as support scales. IN THIS ROLE YOU WILL: Configure and maintain Zendesk and the surrounding support stack, from the day-to-day workflow and automation setup to business rules and permissions, keeping it able to flex and scale as needs change. Build and maintain assignment logic, queues, tagging and taxonomy, and escalation paths, keeping them running cleanly as volume and workflows change. Set up and maintain support tooling, workflows, and access across internal agents, outsourced vendors, and regions, keeping the systems working for each group as the stack changes. Partner with Engineering on the technical requirements for self-service and in-product support surfaces, and build the entry points, help widgets, and routing behind them. Spot the repetitive steps in agent and admin workflows before they become bottlenecks, and build the automations, bulk acti
Higher-paying openings
Jobs with higher listed pay
Staff Software Engineer - Fern
Postman · New York, California, United States
Staff Software Engineer, Business Platform
Postman · San Francisco, California, United States
Principal Software Engineer
Roblox · San Mateo, CA, United States
Principal Software Engineer, Game Safety
Roblox · San Mateo, CA, United States
Staff Software Engineer- Codegen
Postman · Austin, Texas, United States
Sr. Staff Software Engineer, Merchants
Pinterest · San Francisco, CA, US
Related career options
Similar roles with stronger pay
Demand 46/100 · 8 jobs
$840K – $840K/yr
Salary →Demand 43/100 · 5 jobs
$840K – $840K/yr
Salary →Demand 43/100 · 6 jobs
$382.5K – $382.5K/yr
Salary →Demand 43/100 · 8 jobs
$300K – $300K/yr
Salary →Demand 42/100 · 7 jobs
$300K – $300K/yr
Salary →Demand 43/100 · 22 jobs
$278.9K – $278.9K/yr
Salary →Other cities to consider
More places hiring for this role
Get new lead infrastructure software engineer jobs in United States by email
Daily job updates · Unsubscribe anytime