Jobs in United States

Senior Infrastructure Architect in United States

1,941 active opportunities · Updated October 2026

Explore current senior infrastructure architect jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

H
📍 Texas, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$130.7K – $205.2K/yr

Quick readStrong listing-quality and freshness signals

Senior Infrastructure Architect — Enterprise Observability and Automation Description - Job Summary Senior individual contributor responsible for the architecture, implementation, and operational ownership of enterprise observability, monitoring, and automation platforms across HP's global IT environment. This role modernizes infrastructure visibility capabilities while ensuring operational stability, security, and compliance. Serves as a technical and operational bridge between infrastructure engineering, cybersecurity, SOX/compliance stakeholders, automation teams, and external technology partners — leading complex initiatives such as platform migrations, enterprise integrations, and governance enablement. Responsibilities Enterprise Observability and Monitoring Application owner and senior technical authority for enterprise monitoring and logging platforms (Datadog, Splunk), including platform governance, roadmap alignment, and operational oversight. Lead enterprise-scale monitoring platform migrations, including architecture design, agent strategy, data ingestion models, vendor coordination, and deployment across 5,000+ servers. Define standards for alerting, dashboards, observability data quality, and integration with ITSM platforms (ServiceNow). Design and manage multi-org Datadog architecture, including org structure, RBAC, SSO/SAML, secrets management, and cybersecurity compliance. Oversee SNMP-based monitoring of storage and network devices, including device profiling, syslog/event integration, and NetFlow collection. SOX Compliance and IT Governance SOX control owner for enterprise monitoring applications — approve monthly reviews, participate in internal/external audits (EY), and maintain ITGC/SOX compliance. Provide audit evidence, walkthrough docu

AWSAzureAnsible
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions. Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency. About the Role We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments. This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment. The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries. Key Responsibilities Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure. Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios. Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize

PythonAWSGitRest
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $251.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Senior Cloud Security Engineer, you will define and implement the security strategy and controls across our hybrid and multi-cloud environment. Embedded within the Platform Security team, you will operate with a high degree of autonomy, partnering closely with Infosec and Infrastructure teams to secure our cloud infrastructure. Your architectural decisions will directly impact the security of a global platform. You Will: Create Paved Roads: Build innovative services and tooling that make the secure path the easiest path, enabling developers to deploy faster and safer. Engineer Self-Healing Infrastructure: Architect and scale systems that monitor our cloud posture and automatically enforce a self-healing security baseline. Design Secure-by-Default Blueprints: Partner deeply with infrastructure teams to bake threat modeling and secure-by-default patterns into the core DNA of our cloud environments. Implement Effective Guardrails: Deploy organizational security controls that protect our developers without slowing them down. Advance Detections: Write and optimize detections tailored specifically to our footprint. You Have: 4+ years: of relevant professional experience. Experience writing a

PythonAWSKubernetesGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Infrastructure Electrical Engineer, you will help define, validate, and scale the electrical power systems that support high-density AI compute. You will translate evolving compute requirements into practical facility and rack-power architectures, evaluate new technologies and vendor solutions, and drive technical decisions across design, manufacturing validation, construction, commissioning, deployment, and operations. This role is best suited for a senior hands-on engineer with deep experience in mission-critical power systems, strong judgment under ambiguity, and the ability to connect facility infrastructure, hardware requirements, controls, telemetry, reliability, and operations. About the Role We are seeking a senior electrical infrastructure engineer to lead the development of reliable, scalable, and efficient power architectures for high-density, liquid-cooled AI data centers. The ideal candidate has strong practical experience with critical electrical systems at data centers or comparable industrial scale, including medium-voltage and low-voltage distribution, utility interfaces, backup power, UPS and battery systems, rack power delivery, grounding, protection, controls, and monitoring systems. You should be comfortable moving between long-range architecture, detailed engineering review, lab validation, vendor qualification, field deployment, and operational troubleshooting. Key Responsibilities Design and optimize electrical topologies and equipment strategies that reduce cost, accelerate schedules, improve efficiency, increase scalability, and maintain high reliability and maintainability. Review and develop basis-of-des

AWSRestAIGo
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $196.8K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox’s daily active users growing at a record pace, we are seeking a senior data infrastructure engineer to join our new Data Insights team. Our team owns the data tooling that empowers Roblox builders to independently make informed and timely data-driven decisions. As an engineer on the team, you’ll work on the platforms behind tools like Superset, Hex, and Python notebooks, which provide critical insights into the health of our business to users at every level of the company. We tackle diverse challenges in data engineering, infrastructure, and analytics, to deliver the insights our customers need. You will collaborate closely with engineers across our data ecosystem to shape the future of product analytics at Roblox. This role offers the chance to be a founding team member and help define both the technical direction and the long-term shape of the product area from the ground up. This role is a great fit for you if you are proficient in designing and scale robust data infrastructure and applications and have a zeal for developing inspiring, easily maintainable, and reusable code. Join our team and make a significant impact at Roblox. You Will: Architect and deliver a high-pe

TypeScriptPythonReactSQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI is building the infrastructure foundation for the next generation of AI. The Data Center Engineering team defines the strategy, reference architectures, technical requirements, and delivery standards for the large-scale data centers that support OpenAI research, products, and infrastructure partners. As a Data Center Controls Network Engineer, you will design, validate, and scale the controls and OT network architectures that support high-density AI data centers. You will work across controls systems, OT infrastructure, telemetry, commissioning, deployment, and operations, partnering with mechanical, electrical, IT/networking, security, and external delivery teams. About the Role We are seeking a mid to senior OT Network Engineer with a strong controls systems background to lead the design and operation of resilient, secure, and scalable OT network architectures for high-density AI data centers. This role translates compute, power, cooling, and operational requirements into practical OT network designs, evaluates vendor solutions, and drives technical decisions across controls infrastructure, telemetry, commissioning, and operations. The ideal candidate has strong hands-on experience in mission-critical OT environments, including industrial networking, virtualized infrastructure, and OT network operations, with expertise in routing, switching, segmentation, firewall policy, time synchronization, monitoring, and network lifecycle support. Key Responsibilities Define controls, automation, and OT network requirements for AI data center campuses. Develop reference architectures, engineering standards, and reusable design templates. Review and develop basis-of-design and functional design documents, including OT network diagrams, IP/VLAN schemes, telemetry architectures, data flow diagrams, and commissioning requirements. Design OT and infrastructure network architectures, including physical topology, logical topology, IP addressing, subnetting, VLA

PythonSQLPostgreSQLMySQL
H
📍 Louisville, United States
✓ High-confidence listingCompany trend +310%
Quick readStrong listing-quality and freshness signals

Become a part of our caring community *(Selected candidate will be required to live within 60 mins of one of the following metro locations OR be willing to relocate to within 12 months of hire date: Louisville KY, NYC Metro, Dallas Metro, Charlotte NC Metro, Tampa, Miami, Washington DC metro, Chicago, Boston, Atlanta, Nashville) The Senior Security Architect for AI works with EIP Department leaders and Humana enterprise stakeholders to identify, define, and develop security architecture requirements and secure designs for AI technology solutions across Humana's business, information technology, and security domains. The Security Architect leads the development of technical architecture & designs, develop security requirements, perform threat modeling and ensures alignment of security & risk imperatives with business priorities. The Security Architect is responsible for the high-level design and patterns of security program infrastructures to enable the protection of Humana tools, data, systems, and networks. Working with EIP Leaders, the security architect drives alignment between the EIP security strategy, security architecture and infrastructure, and Humana's overall business and technology strategic priorities. In this capacity, the role is responsible for planning, designing, and proposing architectural patterns or security enhancements for Humana's information security and technology infrastructures. The role works with EIP and Humana leaders to review and reconcile Humana business priorities with EIP security requirements. The role engages with relevant EIP stakeholders to identify current and emerging security threats and works to design security architecture elements to mitigate threats as they emerge. Additionally, the role actively works to identify gaps within existing EIP reference architecture and designs updates to impacted sec

Artificial IntelligenceAIRecruitment
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA Architecture Modeling group is looking for Architects, Functional Modeling Engineers, and Simulation experts to join various architecture efforts across GPU/ SOC Architecture teams. A key part of NVIDIA's strength is to innovate in parallel computing fields, delivering the highest performance in the world for high-performance computing. We are constantly looking for ways to improve our SoC and Systems architecture and maintain our leadership. In this position, you will be working with other world-class architects on modeling, analysis and validation of chip & system architectures and features that advance the state of art in performance and efficiency. What you'll be doing: Modeling and analysis of SoC & Systems algorithms and features, across datacenter, automotive, and client products Build and deliver platforms for SOC's that enable left shift for the SW teams aligned with project milestones Work closely with the SOC architects and guide modeling teams to deliver high-quality functional models that involve SOC+GPU use cases Collaborate with our EDA partners to align on customer-facing technologies Develop tests, test plans, and testing infrastructure for new architectures/features and code coverage analysis and reporting Ensure alignment between the various modeling teams at NVIDIA, GPU modeling teams, and modeling teams overseas What we need to see: Master’s or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related relevant field (or equivalent experience) with 5+ years of relevant work experience. Strong programming ability: C++, C along with a good understanding of build systems (CMAKE, make) , toolchains (GCC, MSVC) and libraries (STL, BOOST) Computer Architecture background with experience in modelling wit

PythonDockerAIJenkins
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that our model's training environment is as realistic as possible. This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness. About the Role We're looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You'll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments. This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality. In this role, you will Design, build, and evolve the integration between the Codex harness that powers OpenAI's products and research training infrastructure used for training GPT-next Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance Bu

PythonAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops system and infrastructure solutions designed for the unique demands of advanced AI workloads. We work closely with architecture, infrastructure, and vendor teams to evaluate system performance and guide critical design decisions. Our team focuses on building and applying performance modeling frameworks to understand system behavior, quantify tradeoffs, and support next-generation infrastructure design. About the Role We are seeking an Performance Modeling Engineer to support the development and application of modeling tools used to evaluate AI system performance and inform architectural decisions. In this role, you will partner closely with Senior Performance Modeling Engineers and the Performance Modeling Lead to analyze system behavior, run simulations and analytical models, and help evaluate tradeoffs across compute, memory, networking, and storage. You will contribute to building modeling frameworks while developing a strong foundation in system architecture and AI infrastructure. This role is ideal for early-career engineers with 1–2 years of experience in software engineering, systems analysis, or performance modeling who are excited to grow in large-scale infrastructure and hardware/software systems. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance. Key Responsibilities Support the development and maintenance of performance modeling tools and frameworks Assist in building models to evaluate system behavior across compute, memory, networking, and interconnect subsystems Help analyze distributed system scaling behavior and identify performance bottlenecks Run simulations and analytical models to support architecture and infrastructure decisions Partner with senior engineers to evaluate design tradeoffs across hardware and system components Interpret modeling outputs and help translate findings into clear recommendations Vali

AWSRestAIRust
C
📍 Irving Texas United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

The Applications Development Technology Lead Analyst is a senior level position responsible for building robust, high-performance, large-scale applications. The overall objective of this role is to lead applications systems analysis and programming activities. Responsibilities: * Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements * Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards * Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint * Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation * Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals * Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions * Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary * Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm's reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency. Qualifications: * 5 to 8 years of relevant experience in Apps Development or systems analysis role * Extensive experience system analysis and in programming of software applications * Hands-on experience in

JavaReactDockerKubernetes
C
📍 Irving Texas United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

The Applications Development Technology Lead Analyst is a senior level position responsible for building robust, high-performance, large-scale applications. The overall objective of this role is to lead applications systems analysis and programming activities. Responsibilities: * Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements * Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards * Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint * Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation * Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals * Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions * Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary * Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm's reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency. Qualifications: * 6 to 9 years of relevant experience in Apps Development or systems analysis role * Extensive experience system analysis and in programming of software applications * Hands-on experience in

JavaReactDockerKubernetes
C
📍 Irving Texas United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

The Applications Development Technology Lead Analyst is a senior level position responsible for building robust, high-performance, large-scale applications. The overall objective of this role is to lead applications systems analysis and programming activities. Responsibilities: * Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements * Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards * Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint * Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation * Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals * Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions * Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary * Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm's reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency. Qualifications: * 5-8 years of relevant experience in Apps Development or systems analysis role * Extensive experience system analysis and in programming of software applications * Hands-on experience in We

JavaReactDockerKubernetes
C
📍 Tampa Florida United States, United States
✓ High-confidence listingCompany trend +800%
Quick readStrong listing-quality and freshness signals

The Applications Development Technology Lead Analyst is a senior level position responsible for establishing and implementing new or revised application systems and programs in coordination with the Technology team. The overall objective of this role is to lead applications systems analysis and programming activities. Responsibilities: Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm's reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency. Recommended Qualifications: 6+ years of relevant experience in Apps Development or systems analysis role Extensive exp

SQLGitArtificial IntelligenceAI
N
📍 Santa Clara, United States
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems to solve a complex challenge in accelerated computing: Failure Attribution at Scale. As EDA or equivalent experience workloads scale across thousands of heterogeneous nodes, a single failure can cause massive resource waste. We need an engineer to develop and build an automated framework. This framework will ingest telemetry from CPU and GPU clusters to identify the root cause of job failures in real-time. It will distinguish between hardware faults, infrastructure instability, and software defects. What you'll be doing: Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. What we need to see: Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming. Experience building automated

PythonKubernetesLinuxMachine Learning
🔔

Get new senior infrastructure architect jobs in United States by email

Daily job updates · Unsubscribe anytime