About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem. In partnership with leading cloud providers, hardware manufacturers, utilities, construction partners, and internal engineering organizations, we are delivering hyperscale AI campuses that power the next generation of frontier AI models. Infrastructure Delivery Operations sits at the center of this effort. Our team develops the operating model that connects infrastructure strategy, supply planning, manufacturing operations, and delivery into a single, integrated system that enables OpenAI to deploy AI infrastructure predictably at scale. We partner across Hardware Engineering, Network Engineering, Capacity Delivery, Hardware Operations, Security, Finance, Strategic Sourcing, and external infrastructure partners to create a single, integrated view of program health. Through governance, operational analytics, executive reporting, and scalable operating mechanisms, we enable leaders to proactively manage risk, optimize capacity, and deliver infrastructure predictably at Industrial Compute speed. About the Role We are seeking a Technical Program Manager, Infrastructure Delivery Operations to drive integrated strategy and delivery across OpenAI's rapidly expanding AI infrastructure portfolio. This role sits at the intersection of infrastructure strategy, New Product Introduction (NPI), supply planning, manufacturing operations, and infrastructure delivery. You will lead highly cross-functional programs spanning engineering, supply planning, manufacturing, logistics, construction, commissioning, and operations, ensuring technical and operational dependencies remain synchronized from planning through production readiness. Beyond driving program execution, you will leverage operational insights to improve capacity planning, infrastructure strategy, and deployment readiness. You will also help operationalize new technologies and suppliers by partnering w
Jobs in United States
Infrastructure Engineer in United States
1,504 active opportunities · Updated October 2026
Showing
15 jobs
Explore current infrastructure engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
At Render, we’re building the modern cloud platform for developers creating AI-native, full-stack, multi-service applications. Our mission is to eliminate the tradeoff between the power of hyperscalers and the simplicity of developer-friendly platforms—so teams can ship fast, scale reliably, and focus on their product, not infrastructure. Unlike complex hyperscalers or ephemeral edge/serverless solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability, allowing teams to launch, scale, and manage real-world applications without writing infrastructure code or managing servers. Whether you're building LLM-powered applications, scalable SaaS products, or async processing pipelines, Render empowers teams to move fast and scale confidently from MVP to millions of users. Our platform is trusted by over 6.5 million developers worldwide and continues to grow rapidly. In February 2026, we raised an additional $100M in Series C financing, bringing our total funding to $257M, to accelerate our vision of making cloud infrastructure both powerful and intuitive—designed for the speed of modern AI development. We’re a diverse and talented team that values craft, velocity, and user experience. If you’re excited to help shape the future of the intelligent cloud and empower developers everywhere, we’d love to hear from you. Applying to Render We're seeking candidates who possess high integrity, humility, and an insatiable drive to learn. Through reasoned discussions and continuous feedback, we strive to improve both individually and collectively. We foster an environment of mutual trust and respect, empowering effective debate to achieve the best outcomes for our customers and team. We especially encourage members of underrepresented groups in the tech community to apply and understand that not all successful candidates will meet each requirement listed. Our interview process is unique to each role, an
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re looking for an experienced Executive Assistant to support our CTO (co-founder) and Head of Engineering. This is a highly operational role that goes well beyond calendar management. You’ll own the day-to-day operating rhythm of the Engineering organization, ensuring leaders are prepared, priorities stay coordinated, and critical meetings, communications, and follow-ups happen seamlessly. You’ll partner closely with senior engineering leaders and serve as a trusted point of coordination for employees, customers, candidates, and external partners. Success in this role comes from exceptional organization, judgment, attention to detail, and the ability to keep many moving pieces aligned in a fast-growing environment. RESPONSIBILITIES Own complex calendar management for the CTO and Head of Engineering, balancing shifting priorities while ensuring time is allocated intentionally Ensure leaders are prepared for every day and every meeting by proactively managing agendas, materials, context, logistics, and follow-ups so time is used effectively and decisions move forward Support forward-looking calendar planning, coordinating recurring operating cadences including roadmap planning, leadership meetings, P0 reviews, Engineering All Hands, and other cross-functional forums Own the operational cadence of the Engineering organization, including weekly leadership meetings, monthly Show & Tells, Engineering All Hand
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE At Baseten, we’re looking for a Technical Program Manager to drive our most complex, cross-cutting infrastructure programs. This role will operate across all domains of AI infrastructure, from the GPUs up to the multi-cluster orchestration layer. This is an execution-first role. The work is less about owning a single system and more about imposing order on ambiguity: standing up the right structures, driving decisions to closure, and making sure nothing falls through the cracks across dozens of stakeholders. If you take satisfaction in turning a chaotic, half-defined initiative into a predictable, well-governed program, this role is for you. RESPONSIBILITIES Own complex migrations end to end. Lead large-scale infrastructure migrations across teams and domains. This will involve scoping the work, sequencing dependencies, managing risk, and driving them to completion without surprises. Drive process across infrastructure. Establish and run the operating rhythms that keep programs healthy: planning cadences, status reporting, decision logs, risk reviews, and escalation paths. Make the process light enough that teams adopt it and rigorous enough that it actually works. Help managers build the right structures. Partner with engineering managers and leads to design the team structures, ownership boundaries, and working models a program needs to succeed. Spot gaps in accountability before they become problems. Own fo
At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta's Corporate Engineering team is the infrastructure layer that keeps 1,500+ people connected, secure, and moving fast. This role leads the shift from a reactive support function to a platform function the whole company relies on. Corporate Engineering owns the systems, tooling, and access infrastructure Vanta runs on. From identity and device management to AI tooling infrastructure and ITSM, CE makes sure Vanta's people have the right access, the right tools, and the right guardrails on day one and every day after. CE is at an inflection point. The company has scaled quickly, AI tooling is live company-wide, and the team is ready for a leader who sets technical direction, establishes clear ownership, and builds the leadership layer beneath them. As Sr. Manager, Systems Engineering, you will lead a team of experienced engineers, own several of the company's highest-priority platform initiatives, and make CE a function the rest of Vanta builds on rather than routes around. What you’ll do as a Senior Manager, Systems Engineering at Vanta: Lead, coach, and grow a team of systems engineers and a tech lead. Set clear expectations, hold a high bar, and develop the people who build Vanta's most critical internal systems. Own joiner, mover, and leaver automation end to end in Okta Workflows, so access is correct on day one, follows people through role changes, and closes completely on departure, including API keys and service credentials. Build and own device assurance across macOS and Windows, defining OS, browser, and device-state policy ahead of enforcement rather than in response to it. Rationalize the ITSM estate so self-servi
From $192K/yr
Datadog's Application Performance Monitoring (APM) provides deep visibility into the health, performance, and lifecycle of modern distributed applications, tracing requests from end-user devices (web and mobile) through to backend services. Our goal is to help customers detect root causes faster, optimize application performance, and improve resource efficiency at scale. As the Engineering Manager for APM Serverless, you will help define and deliver the end-to-end serverless APM experience, from auto-instrumentation through troubleshooting, and ensure that OpenTelemetry and Datadog-native customers alike have a frictionless and performant journey. You will also lead efforts to expand coverage of cloud-managed services across providers, ensuring customers can seamlessly trace and monitor critical services in all major and emerging cloud environments. We’re looking for an experienced engineering leader who thrives at the intersection of infrastructure and developer experience. You should care about well-designed APIs, observability-first thinking, and building systems that empower other developers. This is a high-leverage role that will influence how developers across the industry understand and instrument their serverless workloads. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Lead a polyglot team of 8-9 engineers and partner closely with Product and Engineering teams across Datadog to deliver industry-leading serverless capabilities that power consistent, scalable, and intuitive instrumentation across languages. Drive a domain that is technically rich: Lambda, Azure Functions, GCP, OTel billing, Rust, durable functions, distributed tracing across managed services. Engineers on this team work
About the Team OpenAI’s Governance, Risk, and Compliance team helps ensure security and privacy are grounded in how our products and systems actually operate. Assurance Operations partners with Security, Engineering, Infrastructure, Product, Privacy, and Legal to make controls provable, risk decisions explicit, and audit readiness a result of well-designed systems. About the Role We are hiring a technical, product-minded GRC builder who can own consequential audits while improving the control and evidence systems behind them. You will build a reusable common control framework, use Codex to automate assurance work, validate changing system scope, and turn repeated audit friction into measurable improvements. We are looking for someone who questions inherited assumptions, solves novel problems creatively, works closely with engineers, and makes the next audit easier by improving the underlying system. You’ll be responsible for: Lead external, internal, customer, and certification audit work from scoping through evidence review, fieldwork, remediation, and closeout. Build a common control framework linking risk, control intent, implementation, owner, system, environment, evidence, and applicable frameworks. Validate actual scope and ownership instead of assuming last year's controls, product boundaries, or evidence remain accurate. Use Codex to build and test evidence checks, control mappings, request triage, owner workflows, monitoring, and remediation reporting. Partner with engineers on cloud architecture, identity, logging, data flows, software changes, vulnerabilities, and control effectiveness. Design maintainable, permission-aware tools that preserve source provenance, human review, and evidence integrity. Reduce repeated requests and operational burden for control owners through measurable workflow improvements. Define roadmaps, decision rights, milestones, success metrics, and clear cross-functional escalations. We’re looking for someone with: Direct ownership
About the Team Our team turns OpenAI’s latest model capabilities into polished, trusted products for consumers and developers. We build the end-to-end experiences including product surfaces, platform layers, and developer workflows that make cutting-edge AI accessible, useful, and dependable at scale. OpenAI’s Financial Engineering (FinEng) team powers how revenue flows through our products - pricing and packaging, checkout, payments, subscriptions, and the financial infrastructure behind them. We partner closely with Engineering, Data Science, Risk, Finance, and Go-to-Market to make paying for OpenAI products seamless, reliable, and efficient worldwide. We pair rapid innovation with a rigorous approach to responsible deployment. Safety and trust are built into how we design, ship, and learn from real-world usage, so these tools deliver meaningful value while aligning with OpenAI’s mission. About the Role We are seeking an experienced Product Manager to scale the product efforts and technical strategy within our Financial Engineering team. The ideal candidate has prior experience in billing, finance, and accounting, ideally also building solutions for commercial users of varying sizes from small scale to enterprise. This role requires close collaboration with our product, finance, operations, and engineering teams. This position is based in San Francisco, CA. We utilize a hybrid work model with 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Develop a strategy and roadmap to efficiently scale the billing operations and customer experience behind OpenAI’s growing product portfolio Identify and execute opportunities to improve the order-to-cash processing for OpenAI’s largest and most strategic customers Build AI powered tooling for key partner teams such as Finance and User Operations to drive better decisions and business outcomes Collaborate with other product teams to defining OpenAI’s evolving monetization s
About the Role As a Director, Compute & Infrastructure FP&A, you will own and drive the monthly forecasting process for the Compute & Infrastructure org by partnering with various stakeholders across Finance, Accounting, Tax and Engineering. You will play a critical role in planning and forecasting the company’s largest and most complex cost center ( Compute & Infrastructure ). You will collaborate cross-functionally to develop long-range infrastructure investment plans, evaluate build vs. buy decisions, and ensure capital is deployed efficiently to support rapid growth. You will also provide strategic financial guidance through scenario modeling, ROI analysis, and performance tracking, enabling leadership to make high-stakes decisions under uncertainty. What You’ll Do Own compute financial planning & Forecasting. Build and manage consolidation models for GPU/CPU capacity, storage, networking, and data center investments. Translate infrastructure roadmaps into short- and long-term financial forecasts (LRP, annual planning) Coordinate closely with Corporate FP&A on timelines and process Present insights on a monthly basis to senior management. Drive infrastructure investment decisions. Evaluate build vs. buy, vendor vs. owned infrastructure, and capacity allocation tradeoffs. Develop frameworks for investment trade-offs to guide executive decision making. Build scalable tooling & reporting. Implement stakeholder-facing dashboards to track compute spend, utilization, and efficiency metrics. Improve visibility into unit economics (e.g., cost per training run, cost per inference, cost per customer). Drive forecasting accuracy & accountability. Lead budget vs. actual analysis for compute and infrastructure spend. Identify key cost drivers (utilization, pricing, efficiency gains) and reduce forecast variance. Support close & financial reporting. Partner with Accounting to ensure accurate classification of infrastructure spend (OpEx vs C
Leidos has an exciting opportunity for a Sr. DevOps Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary This DevOps Engineer role provides mission critical system support to our customer. You will closely work with the Development team as well as other technology stakeholders to maintain, develop and support IC enterprise products – legacy and new products – in an Agile SAFe environment. The role will also work collaboratively with software engineering to deploy and operate systems. Additionally, this role will help automate and streamline operations and processes; as well as build and maintain tools for deployment, monitoring and operations, and troubleshoot and resolve issues in dev, test, and production environments. Primary Responsibilities: Supports software deployments, cloud infrastructure baselines, and operational availability of production systems. Managing, building, configuring, administering, operating and maintaining all components that comprise the DevOps environment. Defining enterprise Continuous Integration/Continuous Deployment processes and best practices Codifying DevOps best practices across the enterprise Developing and maintaining scripts to automate tool deployment to an AWS cloud environment and other tasks. Scripting and
Leidos has an exciting opportunity for a Sr. Software Engineer in our Intel Security Sector's Analysis Solutions Business Area . Our talented team is at the forefront in Security Engineering, Computer Network Operations (CNO), Mission Software, Analytical Methods and Modeling, Signals Intelligence (SIGINT), and Cryptographic Key Management. At Leidos , we offer competitive benefits , including Paid Time Off, 11 paid Holidays, 401K with a 6% company match and immediate vesting, Flexible Schedules, Discounted Stock Purchase Plans, Technical Upskilling, Education and Training Support, Parental Paid Leave, and much more. Join us and make a difference in National Security! Job Summary As a Software Engineer on this program, you will have the opportunity to build strong systems, software, and cloud environments while providing operations and maintenance for critical systems. This role will provide technical expertise in the design, development, implementation and testing of customer tools and applications. Based in a DevOps framework, this role participates in and/or directs major deliverables of projects through all aspects of the software development lifecycle including scope and work estimation, architecture and design, coding and unit testing. Primary Responsibilities: Participates in and/or directs software programming initiatives using Java, JavaScript, Python, SpringBoot, and Hibernate. Develops software system validation and testing methods using Junit and Katalon and uses integrated custom developed software solutions to leverage automated deployment technologies Develop, prototype and deploy solutions within Commercial Cloud Solutions leveraging infrastructure platform services Coordinate closely with team members, Product Owners and Scrum Masters to ensure User Story alignment and implementation to customer use cases Support th
$115.3K – $149.2K/yr
We’re here for one reason and one reason only – to cure cancer. Every moment is dedicated to developing treatments and every action moves us one step closer to our goal. We’ve made incredible scientific breakthroughs and our pioneering personalized CAR T-cell therapies have changed the paradigm. But we're not finished yet. Join Kite, as we make even bigger advances in cancer therapies, and help shape where our business and medical science goes next. We believe every employee deserves a great leader. People Leaders are the cornerstone to the employee experience at Gilead and Kite. As a people leader now or in the future, you are the key driver in evolving our culture and creating an environment where every employee feels included, developed and empowered to fulfil their aspirations. Join Kite and help create more tomorrows. Job Description Position Summary The Senior IT Quality Engineering Specialist serves as the technical lead for the Kite Laboratory Information Management System (KLIMS) within North America West Coast (El Segundo). This role is responsible for the operational stability, compliance, maintenance, enhancement, and technical governance of the LabVantage platform and associated integrations. The position partners closely with Quality Control, Manufacturing, Validation, Infrastructure, and Kite Business stakeholders to ensure reliable and compliant laboratory operations. Key Responsibilities System Administration & Operational Support • Provide day-to-day administration and technical support for KLIMS and related
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. CVS Health is seeking an SDET Manager to lead Quality Engineering and software delivery transformation across the Medicare Sales & Acquisition ecosystem. This role is critical to modernizing how we build, test, and deliver software, including embedding AI-driven capabilities into the SDLC. You will translate business workflows (Sales, Enrollment, Broker, and Telesales) into scalable, automated testing solutions that improve speed, quality, and operational efficiency. This is a hands-on technical leadership role responsible for driving test automation strategy, enabling intelligent SDLC pipelines, and building a high-performing engineering team aligned to enterprise goals. This role will also help shape the team's transition toward AWS-native test infrastructure and agent-orchestrated cycles, positioning the organization to scale automation alongside delivery volume in future years. Key Responsibilities Lead the design and implementation of AI-augmented Quality Engineering solutions — including agentic test generation, self-healing frameworks, and AI-assisted defect analysis — advancing the SDLC toward an AI-driven Automated Delivery Lifecycle (ADLC) Drive the migration of automation frameworks and test data management onto AWS-native services (e.g., Bedrock for model orchestration, Lambda/Step Functions for event-driven test pipelines), enabling
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. CVS Health is seeking an SDET Manager to lead Quality Engineering and software delivery transformation across the Medicare Sales & Acquisition ecosystem. This role is critical to modernizing how we build, test, and deliver software, including embedding AI-driven capabilities into the SDLC. You will translate business workflows (Sales, Enrollment, Broker, and Telesales) into scalable, automated testing solutions that improve speed, quality, and operational efficiency. This is a hands-on technical leadership role responsible for driving test automation strategy, enabling intelligent SDLC pipelines, and building a high-performing engineering team aligned to enterprise goals. This role will also help shape the team's transition toward AWS-native test infrastructure and agent-orchestrated cycles, positioning the organization to scale automation alongside delivery volume in future years. Key Responsibilities Lead the design and implementation of AI-augmented Quality Engineering solutions — including agentic test generation, self-healing frameworks, and AI-assisted defect analysis — advancing the SDLC toward an AI-driven Automated Delivery Lifecycle (ADLC) Drive the migration of automation frameworks and test data management onto AWS-native services (e.g., Bedrock for model orchestration, Lambda/Step Functions for event-driven test pipelines), enabling
From $154K/yr
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Sr. Staff Software Development Engineer-AI Security to join our team. This is a Hybrid (based in San Jose, CA or Bellevue, WA with a 3 days in office requirement) role, reporting to the Director of Software Engineering in the Emerging Tech department. You will be responsible for designing and implementing core infrastructure components and distributed systems, serving as a foundational architect for our AI security solution. This high-impact role focuses on scaling security infrastructure to support hundreds of millions of users, collaborating with stakeholders across the development lifecycle to drive innovation and technical excellence. What you’ll do (Role Expectations) Architect, develop, and optimize a low-latency, high-throughput AI Security plane utilizing Rust, specifically leveraging its async/await model for highly efficient I/O and service-oriented architecture Build resi
Other cities to consider
More places hiring for this role
Get new infrastructure engineer jobs in United States by email
Daily job updates · Unsubscribe anytime