We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking an accomplished Principal Cloud Storage Engineer to lead the design, engineering, and evolution of our private cloud storage platforms. This role will focus on large-scale storage architecture, data protection, cyber recovery, and resiliency technologies across complex enterprise environments. The ideal candidate will combine deep technical expertise in storage systems with strong leadership, architectural vision, and the ability to influence technical direction across the organization. Key Responsibilities Architect and engineer enterprise storage platforms that ensure data integrity, availability, security, and disaster recovery readiness Design and implement end-to-end storage solutions, including Software Defined Storage, SAN, NAS, and object storage across private cloud and data center environments Drive strategic technology decisions by evaluating emerging products, tools, and standards supporting storage, data protection, cloud, and compute platforms Lead infrastructure initiatives involving storage modernization, data protection, cyber recovery, data migration, and resilience engineering Develop and execute enterprise strategies for backup, recovery, cyber vaulting, and business continuity Create and maintain comprehensive documentation of storage architectures, configurations, policies, and operation
Jobs in United States
Infrastructure Team Manager in United States
1,475 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure team manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
About the Role As Head of Finance - Data Centers, you will be the finance leader for OpenAI’s self-built data center efforts. You will partner with teams across infrastructure, real estate, energy, construction, procurement, and finance to turn proposed sites into sound investment decisions and funded projects. This role spans the full development lifecycle: evaluating opportunities, building investment cases, forecasting capital needs, managing construction budgets, and helping determine how projects should be financed. You will give leadership a clear view of project economics, funding requirements, and risks as we build data center capacity at scale. In this role, you will: Lead financial evaluation of proposed data center and related power infrastructure projects, including site economics, development costs, capacity phasing, lifecycle costs, and key risks. Build and own project-level models and capital expenditure forecasts that connect construction schedules, power delivery, equipment procurement, contingencies, and funding needs. Establish capital budgets and financial controls for active builds. Track commitments, actual spending, change orders, and forecasts to completion; identify cost or schedule risks early. Partner with development, engineering, energy, construction, and procurement leaders on decisions that affect cost, timing, and long-term performance. Work with Treasury, Corporate Finance, Tax, and Legal to evaluate financing options, including project or construction debt, leases, joint ventures, and other partnership structures where appropriate. Prepare investment recommendations and capital approval materials for senior leadership, translating complex project details into clear choices and tradeoffs. Build a consistent portfolio view of project costs, cash requirements, milestones, and financial performance. Partner with Accounting and operations teams through project completion and handoff. You might thrive in this role if you have: Prior exper
From $131K/yr
Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production
About the Role OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models. We’re looking for a Data Scientist to partner closely with Capacity Systems Engineering, Infrastructure, Product, and Research to optimize inference capacity across our global GPU fleet. This role combines statistical modeling, large-scale data analysis, forecasting, and systems thinking to drive critical decisions around infrastructure investments, performance-efficiency trade-offs, and customer experience. You’ll transform complex operational data into actionable insights that directly influence how OpenAI allocates and scales one of the world’s largest AI compute environments. Key Responsibilities Build statistical and machine learning models to profile and improve GPU utilization, latency, throughput, and overall fleet efficiency. Develop forecasting models for inference demand across products, regions, and model families. Analyze production workloads to identify latency bottlenecks and capacity constraints, highlighting optimization opportunities. Partner with Capacity Systems Engineering to inform infrastructure planning and long-term GPU investment strategies. Design experiments and simulations to evaluate scheduling policies, serving strategies, and infrastructure tradeoffs. Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions. Collaborate with Product, Research, Finance, and Infrastructure teams to align compute planning with business growth and model roadmaps. Communicate technical findings clearly to both engineering teams and executive leadership. Qualifications MS or PhD in Statistics, Computer Science, Operations Research, Applied Mathematics, Economics, or related quantitative discipline (or equivalent industry experience). 5+ years of experience working in the infrastructure data science space. Strong ex
From $187K/yr
As a Cloud Security Engineer you will partner with different stakeholders across the organization to secure our cloud infrastructure. As part of the Platform Security organization we secure the building blocks of Datadog’s applications and infrastructure. We do this by building solutions to solve systemic risks and combine an approach of making the secure path easier and the insecure path harder to secure and accelerate the business. We regularly partner with the most bleeding edge internal products and are working to solve and build solutions to enable our safe usage of AI. We also develop AI based solutions to enable security at scale. We are looking for a Service Mesh and Kubernetes focused security specialist to help round out an incredibly strong infrastructure security focused group. You will rotate through a variety of internal projects and gain deep exposure to Datadog’s infrastructure. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Solve our most challenging cloud infrastructure security problems starting with our core building blocks and golden paths. Enable our engineers to build and ship secure solutions quickly. Build and extend Datadog’s Platform Security solutions. Leverage and influence the direction of Datadog’s products to secure our infrastructure, and provide internal feedback that enables our teams to improve the products for ourselves and our customers. Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering or related scientific field or equivalent professional experience. Passionate about advocating for and implementing solutions to complex problems, at-scale, in a large multi-cloud environment. You don’t want to just provide security recommendations, you want to help imple
About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With over 100 enterprise SaaS connectors, flexible LLM choice, and robust APIs, Glean gives organizations the infrastructure to govern, scale, and customize AI across their entire business - without vendor lock-in or costly implementation cycles. At its core, Glean is redefining how enterprises find, use, and act on knowledge. Its Enterprise Graph and Personal Knowledge Graph map the relationships between people, content, and activity, delivering deeply personalized, context-aware responses for every employee. This foundation powers Glean’s agentic capabilities - AI agents that automate real work across teams by accessing the industry’s broadest range of data: enterprise and world, structured and unstructured, historical and real-time. The result: measurable business impact through faster onboarding, hours of productivity gained each week, and smarter, safer decisions at every level. Recognized by Fast Company as one of the World’s Most Innovative Companies (Top 10, 2025), by CNBC’s Disruptor 50, Bloomberg’s AI Startups to Watch (2026), Forbes AI 50, and Gartner’s Tech Innovators in Agentic AI, Glean continues to accelerate its global impact. With customers across 50+ industries and 1,000+ employees in more than 25 countries, we’re helping the world’s largest organizations make every employee AI-fluent, and turning the superintelligent enterprise from concept into reality. If you’re excited to shape how the world works, you’ll help build systems used daily across Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, and many more - deeply embedded where people get things done. You’ll ship agentic capabilities on an open, extensible stack, with the craf
Who we are About Stripe Stripe, LLC. is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterprises to the most ambitious startups - use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career. What you’ll do Responsibilities Drive risk strategy for company-level strategic initiatives, which includes: partnering with the company’s product, legal, operations, sales, and partnerships teams to solve risk problems and ensure a positive user experience. Perform research and analysis to assess Stripe’s current risk performance and develop and prioritize long-term strategic plans for future growth. Increase business enablement, allowing the types of supportable merchants to safely grow at scale, continuously optimizing for efficiency and effectiveness, and stay hyper focused on a positive user experience. Challenge the status quo and provide multiple alternative solutions and key execution criteria. Execute special projects and provide ad hoc analyses as initiatives, products, risks, and opportunities are constantly evolving. Who you are Minimum requirements Must have a Bachelor's Degree or foreign equivalent degree in Law, Business, Policy, Development Studies, or a related field, plus five (5) years post-bachelors, progressive related work experience as a Risk Strategist or a related occupation. Must also have five (5) years of experience in each of the following: Developing risk strategies and solutions and working with technical products, including working with product and engineering, and operations teams for implementation; Distilling complex, ambiguous risk and policy problems into clear guidance for internal sta
Become a part of our caring community The Automation Engineer identifies and implements solutions (hardware and software) for improvement of the high-quality automation infrastructure. The Automation Engineer work assignments are varied and frequently require interpretation and independent determination of the appropriate courses of action. The Automation Engineer designs, programs, simulates, and tests automated processes, and is responsible for detailed design specifications and other documents. Understands department, segment, and organizational strategy and operating objectives, including their linkages to related areas. Makes decisions regarding own work methods, occasionally in ambiguous situations, and requires minimal direction and receives guidance where needed. Follows established guidelines/procedures. Use your skills to make an impact Required Qualifications Bachelor's degree or relevant and equivalent years of experience in lieu of degree requirement. 4+ years of technical experience related to automation. Strong knowledge and understanding of Claude code Experience using AI coding assistants such as Claude code, Github, Copilot, or similar developer productivity tools. Hands-on experience leveraging Claude Code for test automation development, debugging, script generation, and software quality engineering. Experience developing automation using Java, Python, or JavaScript. Experience with Selenium, Playwright, Cypress, or equivalent frameworks. Experience testing REST APIs and backend services. Experience with CI/CD pipelines and automated deployments. 2+ years of experience in Software QA testing in a SAFe Agile environment. Strong experience with black box, web-service integration and server back-end testing.</
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
The DFP Engineer – Manufacturing role defines and implements the validation and screening of new silicon features within high‑volume manufacturing flows. You will translate product requirements into executable test methodologies, infrastructure, and detailed manufacturing test plans that ensure quality, yield, and efficiency at scale. This role sits at the intersection of multiple multi-functional teams to make manufacturing test an outstanding part of the overall codesign and DFP lifecycle. What you will be doing: Own end-to-end manufacturing test methodology across all test stages. Translate system specs and product POR into DFP requirements, test content, coverage, and flows. Define and maintain the DFP roadmap, including infrastructure and turning point planning. Partner multi-functionally to implement test content, debug hooks, and coverage improvements. Drive alignment on manufacturability, test time, binning strategies, and cost vs. coverage trade-offs. Embed testability requirements into design to enable robust screening and debug. Define data and analytics frameworks to support yield analysis and continuous improvement. Lead DFP documentation as the single source of truth and feed findings into future methodologies. What we need to see: MS in Electrical Engineering, Computer Engineering, or related field (or equivalent experience) 6+ years in silicon post‑silicon validation and/or high‑volume manufacturing test for complex SoCs, GPUs, CPUs, or similar. Hands‑on experience with test content bring‑up, limit setting, correlation to characterization, and yield/coverage optimization. Proficiency with scripting and data analysis (e.g., Python, MATLAB, R, SQL) for test data analytics, limit tuning, and yield/debug analysis. <
The Defense Sector at Leidos is seeking a motivated TS/SCI cleared Network Administrator to support the installation, configuration, and day-to-day management of enterprise network infrastructure. This role is an excellent opportunity for an early-career network professional to gain hands-on experience with routing and switching platforms, including Session Smart Router (SSR) / 128 Technology SD-WAN solutions, in a structured and security-conscious environment. The ideal candidate demonstrates a solid foundation in networking fundamentals, a willingness to learn vendor-specific technologies, and the discipline to operate within DoD network standards. The job duties will be performed daily on site at Langley Air Force Base, VA. Roles and Responsibilities: Assist in the configuration, deployment, and ongoing management of routers, switches, and Session Smart Router (SSR) appliances across enterprise and edge network environments. Support the design and implementation of routing policies, service policies, and traffic steering configurations on SSR/128 Technology platforms under senior engineer guidance. Perform LAN switching administration — including VLAN configuration, spanning tree, trunking, and port security — on Juniper EX Series and/or Cisco Catalyst platforms. Assist with the configuration and troubleshooting of routing protocols including OSPF and BGP (eBGP and iBGP) across enterprise WAN and data center environments. Monitor network health, availability, latency, and throughput using network management tools; escalate anomalies and assist in root cause analysis. Support configuration and maintenance of firewall rules, access control lists (ACLs), IPsec VPN tunnels, and other network security controls. Execute software and firmware upgrades, patch management, and lifecycle maintenance activiti
We are seeking a mission-driven Developer Relations Manager focused on Foundational AI Research to engage leading academic labs advancing the next generation of AI models, systems, and methods. In this role, you will work directly with top researchers building frontier AI systems, including large language models, multimodal models, reasoning systems, training methods, inference systems, model serving, and scalable AI infrastructure. You will help researchers adopt NVIDIA’s AI and accelerated computing platforms to push the boundaries of model performance, efficiency, and scale. The ideal candidate brings deep technical credibility in foundational AI, strong research engagement experience, and hands-on expertise in either AI inference research or AI training research. What you'll be doing: Serve as a trusted technical advisor to leading academic AI labs working on foundation models, LLMs, multimodal AI, reasoning, training, inference, and AI systems. Identify high-impact research workloads where NVIDIA software, systems, and accelerated computing platforms can advance model performance, scale, and efficiency. Engage principal investigators, postdocs, graduate researchers, and lab leadership to understand research goals, technical blockers, infrastructure needs, and collaboration opportunities. Track frontier AI research across papers, benchmarks, open-source projects, and academic labs to identify emerging trends and future platform opportunities. Partner with Research Account Managers, Solution Architects, Product, Engineering, and Business Development teams to support researcher adoption and long-term engagement. Represent researcher needs internally by translating academic feedback into actionable insights for product roadmaps, developer programs, education, and platform strategy. Support NVIDIA participation in major AI, ML, and systems research venues through technical content,
About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee
Other cities to consider
More places hiring for this role
Get new infrastructure team manager jobs in United States by email
Daily job updates · Unsubscribe anytime