Jobs in United States

Senior Infrastructure Platform Engineer in United States

1,941 active opportunities · Updated October 2026

Explore current senior infrastructure platform engineer jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

O
📍 Washington, District of Columbia, United States· Full-time
✓ Quality checkedCompany trend -82%

About the Team Join the engineering teams that bring OpenAI’s ideas safely to the world! The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role We’re seeking Software Engineers who can solve complex, high-impact problems across our stack. In this role, you’ll join a nimble team driving the deployment of OpenAI’s technology into new environments and infrastructure that power critical missions in the public sector. You’ll work cross-functionally with product, security, and compliance teams to build the functionality needed to deliver a scalable, reliable platform. You’ll also partner directly with customers to design and build new products and features that create real-world impact. From launching net-new capabilities to optimizing how we serve inference in unique, high-stakes environments, this role offers both breadth and technical depth—giving you the opportunity to shape the future of OpenAI’s technology where it matters most. This role is based in Washington D.C., San Francisco, CA or Seattle, WA. Occasional travel to customer sites is required for this role. In this role, you will: Own the development of new customer-facing ChatGPT and OpenAI API features end-to-end, both on-premises and in the cloud, for our public sector customers. Partner and directly embed with teams across the business, including engineering, security, and compliance, to enable our products to work within the unique constraints of new environments. Talk to users to understand their problems and design solutions to address them Work with the research team to get relevant feedback and iterate on their latest models, developing solutions specific for public sector customers at both the model & data

JavaScriptPythonJavaReact
P
📍 New York, California, United States· Full-time
✓ High-confidence listingCompany trend -100%
Quick readStrong listing-quality and freshness signals

Who Are We? Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman. P.S: We highly recommend reading The "API-First World" graphic novel to understand the bigger picture and our vision at Postman. About Fern, a Postman Company Fern helps software companies build a world-class API experience. Our customers include industry leaders like Nvidia, Square, and Twilio, as well as fast-growing AI companies like ElevenLabs and OpenRouter. In the next year, the majority of API integrations will be implemented by AI agents. Agents don’t read marketing pages. They need strongly typed schemas, structured endpoints, deterministic contracts, and documentation that can be fed into a context window. Our team is small and talent-dense. We’re a team of builders from Google, Palantir, Amazon, and Uber, working together in New York City. About the Role As a Senior Software Engineer at Fern, you’ll build APIs, scale AI infrastructure, and design developer experiences that reach millions of people. Scale infrastructure to keep up with growth: Work on AI systems deployed on Vercel, AWS, and Turbopuffer that must be performant, reliable, and secure under real-world load. Stay close to customers: We work directly with API teams at companies like Nvidia, Twilio, and Square. You’ll debug real production edge cases, shape APIs ba

TypeScriptAWSGitAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxArtificial IntelligenceAI
G
📍 Austin, Texas, United States
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonArtificial IntelligenceAI
C
📍 United States· Remote
✓ High-confidence listingCompany trend +340.2%
Quick readStrong listing-quality and freshness signals

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. POSITION SUMMARY CVS Health is seeking a highly skilled Senior Data Engineer, Observability Engineering to join the Enterprise Observability Platform organization and help advance the next generation of observability, infrastructure, and security data capabilities. The Senior Data Engineer, Observability Engineering will play a critical role in designing, building, and operating scalable data pipelines and data products that power enterprise observability, operational intelligence, and security analytics across the organization. The Senior Data Engineer, Observability Engineering is a senior individual contributor responsible for developing and optimizing Databricks-based data engineering solutions that ingest, transform, govern, and deliver high-volume telemetry, infrastructure, application, and security data. This role combines deep hands-on technical execution with ownership of engineering excellence, operational reliability, performance optimization, and data platform best practices. Working closely with Observability Engineering, Security Engineering, Infrastructure Engineering, and Data Platform teams, the Senior Data Engineer, Observability Engineering will contribute to the evolution of the enterprise observability lakehouse by building resilient ingestion frameworks, establishing data quality standards, enhancing governance controls, and driving efficient, scalable data processing patter

PythonSQLAzure
S
📍 Kalamazoo, Michigan, United States· Remote
✓ High-confidence listingCompany trend +364.7%
Quick readStrong listing-quality and freshness signals

Work Flexibility: Remote As a Senior Lead, Data Engineering, you will serve as a technical leader who helps shape the future of enterprise data solutions. In this role, you will drive complex data initiatives, influence technical strategy, and partner with teams across the organization to build scalable, high-impact data products. This is an opportunity to solve challenging business problems while mentoring fellow engineers and elevating data engineering best practices. What You Will Do Lead the architecture, development, and modernization of scalable enterprise data platforms that support global procurement analytics and business transformation. Define and help execute a multi-year data engineering strategy focused on platform scalability, reliability, automation, technical debt reduction, and long-term maintainability. Design, build, and optimize Azure-based data solutions using technologies such as Databricks, Delta Lake, Azure Data Factory, Azure DevOps, CI/CD pipelines, and infrastructure automation. Integrate and harmonize data across multiple ERP systems by standardizing supplier, purchasing, and master data into common enterprise data models. Partner with procurement analysts, architects, engineers, and business stakeholders to translate complex business needs into reusable, scalable data products and engineering solutions. Establish engineering standards, conduct architecture reviews, improve documentation, and mentor engineers to raise the overall technical capability of the team. Identify and implement AI-enabled approaches that accelerate development, improve data quality, automate documentation, support testing, and enhance analyst productivity. Evaluate and recommend tools, frameworks, patterns, and platform investments that improve performance, reliability, security, governance, and operational ef

PythonReactSQLAzure
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Responsibilities and Duties We are seeking a highly skilled System Tests & Diagnostics Engineer to develop, extend, and integrate specialized silicon validation and diagnostics tools for next-generation AI SoCs. Unlike traditional validation roles focused on executing test plans, this position is responsible for developing the diagnostic software and stress tools that expose hardware failures, characterize silicon behavior, and improve platform observability throughout bring-up and validation. You will work closely with Arm engineers to understand and extend existing diagnostics technologies while developing Graphcore-specific capabilities for future AI hardware. Role Summary You will work with existing Arm-developed diagnostics technologies and extend them to support Graphcore's next-generation AI silicon. You will be responsible for developing system-level diagnostics and stress tools that integrate with an existing framework to detect data integrity, computational correctness, performance, and reliability issues across CPUs, AI accelerators, memory, storage, PCIe, firmware, BMC, and other platform components. Examples include silent data corruption (SDC) tests, power transient stress tools, and platform diagnostics, with opportunities to develop new diagnostics as future hardware capabilities evolve. This role requires close collaboration with hardware architects, firmware enginee

PythonLinuxAIC++
G
📍 Austin, Texas, United States· Full-time
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

About us Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry. As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone. Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation. Job Summary We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments. This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems. The Team The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms. The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems. This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.

PythonAIExcelHR
C-
📍 New York, NY, United States· Full-time
✓ High-confidence listing

$225K – $300K/yr

Quick readStrong listing-quality and freshness signals

CLEAR is building THE secure identity company of the future. Our mission is to make experiences safer and easier—physically and digitally. With more than 43 million Members and a growing network of partners across the world, CLEAR's secure identity platform is transforming the way people live, work, and travel. Whether it’s at the airport, stadium, or throughout your everyday life, CLEAR unlocks the magic of frictionless experiences. As a Senior Software Engineer, Data, you will design, build, and operate the next generation of our data platform and products – going beyond ID to power a networked digital identity – while keeping member privacy, security, and reliability at the core. What you’ll do: Build and operate scalable, reliable data systems and pipelines – from ingestion to modeling to visualization – so Analysts and Engineers can self-service changes in an automated, tested, secure, and high-quality manner. Develop and maintain end-to-end data products and pipelines (batch and/or streaming) that collect, clean, transform, and model data, and own the infrastructure that powers them to unlock new business use cases and reporting. Implement and maintain infrastructure-as-code, CI/CD, and shared developer tooling for data products (e.g., Pulumi/Terraform, GitHub, orchestration tools like Dagster/Airflow) to make it easy and safe for teams to build, test, and ship changes across environments. Improve the security, compliance, and cost posture of the data stack through robust dependency management, IAM and secrets hardening, observability, and performance/cost optimizations. Partner with product and other stakeholders to uncover requirements, make architectural decisions, and continuously improve our data platform and processes. How you’ll measure success: Data reliability & SLAs: % successful pipeline runs, adherence to freshness SLAs for core datasets, and reduction in data-related incidents impacting stakeholders. Platform quality & efficiency: Reductio

PythonSQLAWSCI/CD
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend +66.7%

From $224K/yr

Quick readStrong listing-quality and freshness signals

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Solutions Engineer (Enterprise Pre-Sales) Secure Every Identity, from Human to AI Agent About the Role At Okta, we believe that identity is the foundation of security and digital transformation. As a Senior Solutions Engineer, you will be the trusted technical advisor to our largest, most complex enterprise customers and a critical driver of our sales organization. We are looking for a highly strategic Pre-Sales professional whose consultative expertise, emotional intelligence, and enterprise sales acumen are their defining strengths. While technical agility is required, your primary focus will be owning the technical sales cycle by anchoring technical features to positive business outcomes. You will partner closely with Enterprise Account Executives to uncover top-of-mind business challenges—specifically around mitigating risk, reducing costs, and driving operational efficiency. By establishing value-based conversations, you will prove how Okta’s independent, neutral, and end-to-end identity platform can transform their architecture and secure the "Tech Win" on large-scale deals. What You'll Be Doing Master the Discovery Process: Leverage exceptional active listening to dig deep into customer pain points. You will uncover the 'why' behind the initiative, focusing on how Okta's vast pre-built integrations and scalable architecture can solve their most complex, diverse environmental challenges. Navigate Complex Organizations: Translate highly technica

AWSGitRestMachine Learning
V
📍 United States· Full-time· Remote
✓ High-confidence listingCompany trend -90%
Quick readStrong listing-quality and freshness signals

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Our Senior Software Engineers lead and mentor engineers, delivering high-value products for our customers and infrastructure that enables our business to scale. As a Senior Software Engineer, you'll be responsible for setting technical direction to enable our product and infrastructure to scale with our business, driving complex projects across our technical stack, and mentoring our talented engineering team. Your past experience will be leveraged to enable and accelerate Vanta's growth. Our business has found incredible product-market fit and has monetized effectively since the day we signed our first customer. We're growing at a blistering pace, which presents career-defining opportunities for engineers to accelerate their growth and to contribute to a rapidly-scaling company. Visit our Vanta Engineering Blog to learn more about what our team is working on! The Authoring Experience team owns how Vanta tests are authored and deployed to customers, from the APIs and services underneath to the experiences built on top. As a Senior Backend Engineer on this team, you will own the services and public-facing customer APIs that turn Vanta's automation platform into products customers touch directly, stitching together systems across the platform to serve them well. What you’ll do as a Senior Fullstack Software Engineer at Vanta: Own the architecture and rollout strategy for critical platform services, from initial design through large-scale rollout and migration. Set the technical bar for how we design, publish, and maintain customer-facing APIs while owning things like security, reliability, and developer experience. Stitch together

TypeScriptReactNode.jsMongoDB
H
📍 New York, NY, United States
✓ Quality checkedCompany trend +310%

Become a part of our caring community Most AI engineering jobs are a thin wrapper around a model API. This role is different. We build the platform that transforms millions of clinical documents into trusted, actionable data. Our systems use large language models (LLMs) to read medical records, extract structured facts, answer complex questions with citations back to the source document, and route ambiguous cases to human experts for review. Our users make decisions that impact real healthcare outcomes, so “good enough” is not good enough. Building AI systems that are accurate, reliable, auditable, and scalable is at the core of this role. As a Senior AI Applied Engineer, you will design, build, deploy, and operate production AI systems used at scale within one of the largest health insurers in the United States. You will own solutions end-to-end, from user experience and APIs to model orchestration, evaluation frameworks, infrastructure, and production operations. Why Join Us Build production AI systems where LLMs are in the critical path, not just demos or proofs of concept. Work on extraction, retrieval, agentic workflows, and human-review systems that process real healthcare data at scale. Own projects end-to-end across frontend, backend, AI orchestration, infrastructure, deployment, and operations. Solve challenging problems around accuracy, explainability, traceability, and reliability in regulated environments. Ship quickly in a small, high-impact team that embraces AI-assisted development and rigorous quality standards. Build systems that continuously improve through expert feedback, evaluations, and human-in-the-loop workflows. Key Responsibilities Design, develop, and deploy full-stack AI-powered application

JavaScriptTypeScriptPythonReact
V
📍 United States· Full-time
✓ Quality checkedCompany trend -90%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Senior Software Engineer, Trust (TPRM) - Job Description As a Senior Software Engineer on Vanta's Trust TPRM team, you'll build the full-stack product experiences and underlying data infrastructure that help enterprises manage vendor risk at scale — working across teams focused on vendor lifecycle management and vendor monitoring. Vanta's Trust TPRM (Third-Party Risk Management) team is building the products that make vendor risk management seamless for security and procurement teams. From vendor onboarding and lifecycle management to continuous monitoring and procurement integrations, we're creating the platform that helps Vanta customers understand, track, and mitigate third-party risk — a fast-growing, business-critical capability for modern enterprises. As a Senior Software Engineer, you'll contribute as a core member of either the Vendor Lifecycle or Vendor Monitoring Experience team. You'll design and ship full-stack features that directly shape how customers manage vendor relationships, collaborate with product and design partners, and bring real engineering ownership to a product area that's growing quickly within Vanta. What you’ll do as a Senior Software Engineer at Vanta: Design, build, and maintain full-stack features across the TPRM product surface, including vendor onboarding, lifecycle management, and monitoring workflows Contribute to the vendor data model and core platform abstractions that power TPRM products Write clean, well-tested code and actively participate in code reviews; uphold engineering quality standards Engage in architecture discussions and contribute to technical decision-making within your team

TypeScriptReactNode.jsRest
C
📍 United States· Full-time
✓ Quality checkedCompany trend -100%

At ClickUp, we're building the future of work: the first truly converged AI workspace unifying tasks, docs, chat, calendar, and enterprise search, all supercharged by context-driven AI. We are an AI-native company. Every team member is expected to leverage AI daily, and we evaluate AI fluency as part of our hiring process. Join us and help redefine what's possible. 🚀 The Mission The Foundry is ClickUp's internal AI innovation lab — embedded inside GTM Systems and accountable for turning AI capabilities into production-grade, internally deployed products that make every GTM function faster and smarter. We build the infrastructure that powers AI-first work across Sales, Marketing, Post-Sales, and Revenue Operations. As the Senior Software Engineer on this team you will own the technical delivery of our MCP server platform, agent orchestration layer, and internal tooling — shipping production systems used daily by hundreds of ClickUp employees, and scaling your own throughput by treating AI tools as first-class engineering collaborators. What You'll Own MCP Server Platform Design, build, and operate Model Context Protocol servers that expose CRM, ticketing, analytics, and communication data to AI agents across the GTM stack Implement Okta PKCE authentication flows and RBAC policy enforcement so agents access only the data they're authorized to touch Maintain deployment infrastructure on AWS (Bedrock, Lambda, ECS, API Gateway) and contribute to GCP workloads where applicable Own observability: structured logging, distributed tracing, latency SLOs, and on-call runbooks for every production server Agent Orchestration & AI-Native Products Build and maintain multi-step autonomous agents that execute end-to-end GTM workflows — lead qualification, deal room assembly, onboarding automation, support triage, and more Architect prompt engineering frameworks, tool-call schemas, and agent evaluation harnesses that make AI behavior predictable and auditable Integrate with LLM p

TypeScriptPythonReactNode.js
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $525.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Senior Director, Generative AI About the Role Roblox Build is our generative creation product, the platform where creators design, build, and publish 3D experiences. We are looking for a Senior Director of Generative AI to lead the Applied AI organization inside Build, responsible for turning state-of-the-art foundation models into high-quality, reliable creation systems at Roblox scale. This leader will own the full applied AI stack: model strategy and routing, model adaptation and fine-tuning, code generation (CodeGen), 3D layout generation (LayoutGen), and the evaluation science and infrastructure that tells us what actually works. You Will Own model strategy and routing for Build. Design and build an intelligent model layer that selects the right model for each creation task based on quality, capability, latency, cost, and safety, leveraging both frontier models and Roblox-adapted open-source models. Lead model adaptation across the Applied AI org, including fine-tuning, distillation, synthetic data generation, human feedback pipelines, and preference optimization for Roblox-specific creation tasks such as Luau code generation and 3D scene understanding. Drive CodeGen capabilities

AWSGitRestMachine Learning
🔔

Get new senior infrastructure platform engineer jobs in United States by email

Daily job updates · Unsubscribe anytime