Jobs in United States

Senior Infrastructure Platform Engineer in United States

1,941 active opportunities · Updated October 2026

Explore current senior infrastructure platform engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $326.1K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Security Software Engineer on the Production IAM team, you will set the technical direction for how identity and access work across Roblox's production infrastructure, from the mTLS-based identity that services use to authenticate to one another, to the privileged access controls that govern how engineers reach production. The team is accountable for Roblox's machine and workload identity platform, its centralized authorization engine, its production access management platform, production PKI and certificate lifecycle, and just-in-time privileged access for engineers. As an individual contributor in Production IAM, you will define multi-year strategy, drive alignment across Roblox Platform, mentor senior and staff engineers, and personally build the hardest parts of these systems. As AI agents become first-class actors in production, you will also help pioneer how they get identity, prove who they are, and receive safely-scoped access. You will Lead the architecture for production identity and access. Define and evolve the end-to-end design for machine, workload, human, and AI-agent identity across our hybrid on-prem and cloud fleet, making secure access invisible when

PythonJavaAWSGit
H
📍 Texas, United States of America, United States
✓ High-confidence listingCompany trend +103.7%

$105.1K – $161.8K/yr

Quick readStrong listing-quality and freshness signals

Senior React Native Mobile and Full Stack Developer Description - We are seeking a highly skilled Senior React Native Mobile and Full Stack Developer to design, develop, and maintain secure, scalable, enterprise-grade mobile applications for iOS and Android , as well as the cloud services and APIs that support them. The ideal candidate brings extensive experience building production-ready cross-platform mobile applications with React Native , combined with strong backend and cloud development expertise using object-oriented programming languages such as C#, Java, Python, or Go to deliver secure, end-to-end solutions. This role requires deep expertise in mobile authentication , secure storage , offline-first architecture , data synchronization , push notifications , GPS/location services , and mobile security best practices , as well as experience designing and developing RESTful APIs , integrating with cloud platforms such as Microsoft Azure, or AWS , and building scalable backend services. You will work closely with Product Managers, UX Designers, Security Architects, Backend Engineers, and QA teams to deliver high-quality, mission-critical mobile solutions used by enterprise customers worldwide. The ideal candidate is a hands-on specialist who is passionate about building high-quality software, driving architectural decisions, mentoring engineers, and delivering reliable, secure, and performant mobile experiences across the full technology stack. Key Responsibilities Design, develop, and maintain secure, scalable, and high-performance cross-platform mobile applications using React Native for both iOS and Android . Design, develop, and maintain backend APIs , cloud services, and supporting infrastructure that

PythonJavaReactAWS
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product workloads. The company seeks a Senior Technical Program Manager (TPM) to lead complex, cross-functional programs powering NVIDIA’s next-generation AI software platforms. In this role, you will drive software initiatives across platform services, cloud infrastructure, and system integration. The focus is on enabling scalable, reliable, and supportable software for AI workloads. You will be responsible for managing high-impact engineering programs within a dynamic, fast-paced roadmap, aligning priorities across teams, and ensuring timely, high-quality delivery. This role requires strong technical competence, a proactive approach, and the ability to operate effectively across multiple levels of the organization. This is a software-first TPM role. The ideal candidate has extensive experience managing software initiatives. They also understand the full-stack environment, including infrastructure dependencies, system bring-up, integration readiness, and operational needs to support software across stack layers. What You'll Be Doing: Lead end-to-end execution of software platform initiatives, including planning, execution, delivery, and operationalization. Work together with software, infrastructure, product, and operations teams to ensure alignment on goals, deliverables, achievements, and schedules. Lead cross-functional initiatives encompassing cloud-native services, platform software, system integration, and release delivery. Help connect software roadmap execution to full-stack readiness, including dependencies across infrastructure, bring-up, validation, and downstream operational support. Identify cross-functional dependencies, mitigate risks, and drive resolution of complex technical and programmatic issues. Establish clear success metrics and reporting mechanis

KubernetesLinuxMachine LearningArtificial Intelligence
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -8%

NVIDIA is seeking a Senior Technical Program Manager to join the CSP Engagements team, focused on deep technical engagement with hyperscale cloud service providers for NVIDIA’s next‑generation datacenter systems such as Vera Rubin NVL72. This role is intended for experienced systems and embedded software leaders—including software engineering managers, technical leads, or senior architects—who have led datacenter server and platform software programs and can operate as a trusted technical partner to hyperscale CSP engineering teams. As a member of the CSP Engagements team, you will act as the primary technical engagement leader between NVIDIA’s system software organizations and CSP platform, system software, and AI teams, ensuring alignment, readiness, and successful large‑scale deployment of NVIDIA‑based datacenter solutions. What you will be doing: Lead deep technical engagements with hyperscale CSPs as the primary NVIDIA point of contact for system software, firmware, and platform readiness for NVIDIA datacenter products. Partner directly with CSP system software, firmware, and infrastructure engineering leaders to align on software architecture, bring‑up plans, deployment readiness, and production requirements for NVIDIA‑based server and rack‑scale platforms. Represent CSP technical priorities internally, advocating for customer requirements and tradeoffs across NVIDIA’s system software, firmware, hardware, silicon, and product teams are aligned to customer needs, timelines, and constraints. Own the end‑to‑end CSP engagement lifecycle, from early technical alignment and pre‑production readiness through large‑scale deployment, escalation management, and sustained production support. Drive bi‑directional technical communication: translating CSP system‑level requirements into actionable focus areas for NVIDIA engineering teams, while clearly communicating N

LinuxArtificial IntelligenceAI
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $226.4K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. As Senior Technical Program Manager for the Human-in-the-Loop (HITL) Operations , you will own the end-to-end lifecycle of Roblox's AI data labeling and model evaluation ecosystem. This is a high-leverage, high-ownership role at the intersection of Research, Engineering, and Operations. You will scale our data operations that is handling million of items evaluated annually today with projection of 4x the current volume by 2030, managing a multi-million annual budget and a distributed workforce of 100+ remote contractors — all while driving the platform evolution from manual workflows to AI-assisted, LLM-accelerated operations. This is a rare opportunity to build the data infrastructure that directly determines the quality of Roblox's AI models across creator tools, content discovery, in-experience AI, and 3D generative content — at the exact moment when human judgment is the critical ingredient for getting these models right. You Will Lead data programs end-to-end. Own the full lifecycle of labeling and model evaluation workflows across Roblox's AI teams — from translating ambiguous ML requirements into structured annotation tasks, to overseeing contractor execution, quality review, a

PythonSQLAWSGit
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $192K/yr

Quick readStrong listing-quality and freshness signals

As the Senior Product Manager for the Actions & Automations team, you will own the ecosystem that enables customers, partners, and Datadog teams to build, deploy, and operate AI agents on Datadog. You will drive the strategy and execution for the platform capabilities, developer experience, integrations, and extensibility model that make Datadog the best place to build agents that understand and act on production systems. Modern engineering organizations are entering a new era where software is not only monitored and operated by humans, but increasingly by AI-powered agents. As agentic workflows reshape how teams build, operate, secure, and troubleshoot systems, customers need a platform for creating specialized agents, connecting them to business and engineering systems, governing their behavior, and extending them to solve unique organizational problems. You will define and build the ecosystem that makes this possible. Agent Builder sits at the intersection of Datadog's products, AI capabilities, and ecosystem strategy. You will have the opportunity to work across the breadth of the Datadog platform, partner with teams throughout the company, and help establish Datadog as the foundation for operational AI. At Datadog, we place value in our office culture, the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Define the vision, strategy, and roadmap for Datadog's Agent Builder platform and ecosystem. Own the core platform capabilities that enable customers and partners to create, customize, deploy, and manage AI agents. Drive the extensibility model for agents, including integrations, tools, actions, context sources, APIs, SDKs, and developer workflows. Shape how agents perform actions across Datadog products and third-party systems. Partner closely with AI, platform, infrastructure, and product tea

AIGoRustSpring
Z
📍 Bellevue, Washington, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler office; remote can be considered for exceptional candidates) reporting to the Senior Manager, Site Reliability Engineering in the Zero Trust Exchange department. As a key member of the Zero Trust Exchange team, you will own the systems-level reliability and performance of Zscaler’s high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions across a global, multi-region fleet. This is a software-first SRE role: you will write production-grade code and automation, drive the shift from reactive incident response, and bring engineering discipline to the systems-level work - OS, network and application debugging - that keeps the fleet operating safely at scale. What You’ll Do (Role Expectations) Maintain h

PythonKubernetesLinuxAI
Z
📍 Bellevue, Washington, United States· Full-time
✓ High-confidence listing

From $180K/yr

Quick readStrong listing-quality and freshness signals

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location. Distributed across 160+ public exchanges globally and thousands of private exchanges at the edge, the SASE-based Zero Trust Exchange is the world’s largest in-line cloud security platform. We believe the future of work is Human + AI and are building an AI-native enterprise where human potential is amplified by machine intelligence to solve the world’s hardest security challenges. Driven by deep customer obsession, we are committed to the mission, outcome, and to each other. We bring these commitments to life through three core behaviors: ownership and collaboration, trust through outcomes and impact, and a challenge culture with ongoing feedback. Ready to make an impact at the company pioneering security transformation in the AI era? Join us at Zscaler. Role We are looking for a Senior Staff Rust Developer to join our Platform Convergence Team. This is a hybrid role based in San Jose, CA reporting to the Sr. Director, Software Engineering. Join us to build a new platform from the ground up that can scale hundreds of millions of users with high reliability and low latency. You will design and implement distributed system and core infrastructure components while collaborating closely with various stakeholders. What you’ll do (Role Expectations) Design and build a low-latency, high-throughput data forwarding plane using Rust, leveraging its async/await model for efficient I/O and service-oriented infrastructure Develop distributed, scalable systems with a focus on concurrency, fault tolerance, and messaging Implement and maintain gRPC-based APIs and services to integrate forwarding plane capabilities with control and orchestration layers Optimize system

AWSKubernetesCI/CDGit
C-
📍 New York, New York, United States· Full-time
✓ High-confidence listing

$145K – $170K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. CLEAR is seeking a Senior Security Operations Analyst III to join our SOC team to help strengthen our ability to detect, investigate, and respond to evolving security threats. In this role, you’ll lead complex investigations, improve CLEAR’s threat detection and response capabilities, and serve as a trusted security partner while helping develop the analysts and program around you. What you'll do: Lead complex investigations of security events across corporate networks, endpoints, data centers, cloud environments, and other critical systems, driving incidents from initial analysis through escalation and remediation Develop, tune, and optimize threat detection logic across SIEM, EDR, and other security platforms, proactively identifying coverage gaps, reducing false positives, and improving the fidelity of security alerts Partner with Engineering, Infrastructure, and other teams to investigate threats, identify root causes, communicate risk, and drive timely remediation and improvements to CLEAR’s security posture Apply threat intelligence, data, automation, and AI-enabled tools to identify emerging attack patterns, accelerate investigations, improve detection workflows, and strengthen decision-making while applying sound security judgment Serve as a subject matter expert and escalation point for other analysts, mentoring junior team members, sharing knowledge, and helping establish scalable processes, playbooks, and standards for threat detection and analysis Continuously evaluate CLEAR’s detection coverage against the evolving t

GitRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -82%
Quick readStrong listing-quality and freshness signals

About the Team The Core Services organization builds and runs the mission-critical online services that product teams rely on in production. We own foundational distributed systems and platform capabilities that enable reliable execution, high-performance services, and large-scale file/data needs across our products. This team is distinct from developer infrastructure and data infrastructure—our focus is production service foundations and core runtime services. About the Role We’re hiring an Engineering Manager, Core Services to help lead teams responsible for highly reliable, high-scale distributed systems that sit on the critical path for OpenAI products. Your team will own foundational production systems that OpenAI’s product engineering teams build on. You’ll collaborate closely with product and infrastructure partners to ship reliable services quickly, and help scale systems and teams as OpenAI grows. You’ll partner closely with senior engineering leaders to scale the org, mature operations, and drive major platform initiatives. This role requires strong technical ability. You’ll be responsible for: Managing and growing a high-performing team of infrastructure engineers. Leading teams building and operating large, critical production platforms, including cluster reliability, scaling, and rollout safety. Building and operating mission-critical distributed systems with strong operational rigor (SLOs, incident response, capacity planning, reliability). Setting technical direction for platform foundations such as workflow/orchestration capabilities, large-scale file/blob/storage services, and core service foundations. Partnering with a broad set of stakeholders, including product engineering, adjacent infrastructure teams, and (where relevant) finance/cost partners. Coaching, mentoring, and developing engineers and emerging leaders. You might thrive in this role if you: Have significant experience leading teams that run mission-critical infrastructure in production

AWSRestAIGo
D
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -85.2%

From $220K/yr

Quick readStrong listing-quality and freshness signals

The Applied AI team designs and builds algorithmically driven features in the Datadog app. We work across a range of applications, primarily focusing on analysis on streaming data such as anomaly detection, error outliers and faulty deployment analysis. As an Applied Scientist you will work on building models and algorithms for machine learning powered features within the Datadog platform. You will work closely with our engineering and product partners to explore, build, scale and deliver these features that we incubate within the Applied AI team. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Design solutions for our different use cases. You will research and benchmark relevant algorithms to find the best fit for our use-cases Leverage machine learning algorithms and statistical techniques to build new scalable product features Develop, deploy and monitor new and existing features to production Participate in our journal club by reading and presenting the latest academic research papers to the team Explore, analyze and tell the story behind high volumes of data flowing through Datadog systems Maintain and monitor the models, services and infrastructure owned by your team Participate in your team’s on-call rotation Who You Are: You have a BS/MS/PhD in a Computer Science, Engineering, Machine Learning or related scientific field or equivalent experience You have experience working with high-scale systems and datasets including building models, applying machine learning to real business problems, and writing production data pipelines You can explain complex ideas and algorithms to non-technical audiences You care about code simplicity and performa

Machine LearningAIGoRust
G
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

👋 Welcome to Glide! At Glide we’re reimagining the banking experience for the modern world . Our embedded fintech platform empowers legacy financial institutions, like community banks and credit unions, to pioneer novel digital experiences for their customers. You’ll be joining an all-star team with engineering, product, and growth experience from Stripe, Google, and Amazon. We’re looking for a talented early career Fullstack Engineer to help us build our platform. We’re bringing a new perspective to the decades-old financial world , and we’re hoping you can help us do that! Why You’ll Love This Role (Especially if you’re a new grad): A unique opportunity to work side-by-side with Glide’s founders and senior engineers, learning directly from industry veterans. Gain hands-on, end-to-end experience across the full stack: from architecture design to deployment. Be part of a small, fast-moving startup, where your code has immediate impact. Get mentorship and exposure to real-world engineering best practices, product development, and startup culture. Your Responsibilities: Move seamlessly between frontend and backend to build a well-tested, secure frontend for our core web product. Create trustworthy, safe user experiences by building interfaces that are simple, reliable, and performant using tools like Typescript, React, Node.js and NextJS . Design a scalable architecture that can serve hundreds of thousands of end users. Bring designs to life through beautifully-crafted code. Articulate a long-term technical direction and vision for maintaining and scaling our web product suite. Lead frontend and backend infrastructure & tooling for an ambitious product roadmap. Need-to-Haves: Experience with Javascript and tools like Typescript, React, Node.js and NextJS . Experience with modern, responsive HTML & CSS . Experience using data fetching libraries like React Query/Tanstack or tRPC to synchronize client and server data. Excellent understanding of software engineer

JavaScriptTypeScriptJavaReact
G
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

👋 Welcome to Glide! At Glide we’re reimagining the banking experience for the modern world . Our embedded fintech platform empowers legacy financial institutions, like community banks and credit unions, to pioneer novel digital experiences for their customers. You’ll be joining an all-star team with engineering, product, and growth experience from Stripe, Google, and Amazon. We’re looking for a talented early career Fullstack Engineer to help us build our platform. We’re bringing a new perspective to the decades-old financial world , and we’re hoping you can help us do that! Why You’ll Love This Role (Especially if you’re a new grad): A unique opportunity to work side-by-side with Glide’s founders and senior engineers, learning directly from industry veterans. Gain hands-on, end-to-end experience across the full stack: from architecture design to deployment. Be part of a small, fast-moving startup, where your code has immediate impact. Get mentorship and exposure to real-world engineering best practices, product development, and startup culture. Your Responsibilities: Move seamlessly between frontend and backend to build a well-tested, secure frontend for our core web product. Create trustworthy, safe user experiences by building interfaces that are simple, reliable, and performant using tools like Typescript, React, Node.js and NextJS . Design a scalable architecture that can serve hundreds of thousands of end users. Bring designs to life through beautifully-crafted code. Articulate a long-term technical direction and vision for maintaining and scaling our web product suite. Lead frontend and backend infrastructure & tooling for an ambitious product roadmap. Need-to-Haves: Experience with Javascript and tools like Typescript, React, Node.js and NextJS . Experience with modern, responsive HTML & CSS . Experience using data fetching libraries like React Query/Tanstack or tRPC to synchronize client and server data. Excellent understanding of software engineer

JavaScriptTypeScriptJavaReact
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior Android engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable Android foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Have a proven track record of building high-quality Android applications in production. Are fluent in Kotlin (and/or Java) and familiar with Android development tools and architecture components. Prioritize performance, security, and user experience in mobile development. Enjoy working cross-functionally to bring ambitious product ideas to life. Care deeply about performance, security, and user experience. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We

JavaAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior iOS engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable iOS foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. In this role, you will: Build and ship new experiences on iOS that showcase the power of AI. Optimize app performance, reliability, and responsiveness at global scale. Design and maintain shared iOS frameworks and primitives for account, trust, and commerce flows that are used across OpenAI’s mobile apps. Establish robust testing frameworks and refine app architecture for long-term maintainability. Collaborate with product, design, research, and backend teams to deliver high-impact features. Provide technical leadership to shape the future of OpenAI’s iOS platform. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Hav

AWSRestAISwift
🔔

Get new senior infrastructure platform engineer jobs in United States by email

Daily job updates · Unsubscribe anytime