We're building the platform that lets long-running autonomous agents operate safely inside NVIDIA's enterprise. These are not assistants on a developer's laptop. They are fleets of agents deployed in the cloud, running continuously at scale on shared accelerated compute. They take on real work across enterprise systems, so people get far more done than they could before. This role defines the constructs that agents are built from: the blueprints they start from, the tools, skills, and plugins that power them against enterprise data, the runtime safety harness that keeps them in bounds, and the connections into credential management, sandbox, memory, and observability. The team designs and ships these building blocks so that agent developers across the company can stand up a new agent, wire it in, and run it for days or weeks. Security and safe execution come out of the box, not something each team has to get right on its own. Today an agent runs inside a single harness. Claude, Codex, and open-source agent harnesses each work differently underneath, with their own execution model, tool interface, and telemetry shape. The platform smooths over those differences, so a single skill, safety policy, or trace works the same no matter which harness is running. We want to enable agents that act on a person's behalf, governed and secure, continuously evaluated and self-improving. These agents coordinate and hand work off to each other, with identity and policy following every hop. They route and tune themselves across harnesses from live eval signals, and get better from their own production telemetry instead of waiting on a human to retrain them. Have you run agents on a harness like Claude or Codex and hit the walls that show up when they run for real, for days, against live systems — and wanted them to learn from it on their own? We're building the platform that solves those problems once, for every team. What you'll b
Jobs in United States
Platform Architecture Manager in United States
3,643 active opportunities · Updated October 2026
Showing
15 jobs
Explore current platform architecture manager jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Health100 is an AI‑native health technology platform that unifies pharmacies, providers, insurers, PBMs, and digital health solutions into a single, consumer‑focused ecosystem. Powered by Google Cloud AI, we’re reimagining personalized and connected health experiences. As a Senior Software Engineer for Health100, you will play a crucial role within a collaborative team — designing, developing, and maintaining backend services and APIs while ensuring releases are well-coordinated, fully prepared, and successfully deployed to production. The ideal candidate brings strong technical expertise in modern backend development, excellent problem-solving skills, and a proactive approach to production monitoring, issue triage, and cross-team coordination. This position is critical in maintaining high engineering standards, ensuring smooth release cycles, and driving operational excellence across the development lifecycle. *This role can be based anywhere in the US; hybrid or remote with preference for candidates to work out of our corporate headquarters in Woonsocket, RI. Responsibilities: Partner with technical leaders and the open-source community to contribute to technical designs, frameworks, roadmap definition, and requirements-gathering. Provide domain knowledge and engineering insight to guide early designs, ac
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary Health100 is America's trusted front door to health and care. The Health100 platform integrates any participating health plan, PBM, pharmacy (retail and specialty), provider, digital health point solution provider, and employer, and addresses the top health care challenges for the consumer. We are looking for a hands-on, passionate engineering leader to join a high-energy, mission-driven team on the forefront of digital health innovation — reinventing how consumers engage with their health. As the Lead Director of Software Engineering, you will lead the strategic vision, development, and execution of our platform engineering roadmap. You will drive innovation, manage cross-functional teams, and ensure the delivery of high-quality products that exceed consumer expectations. As a key strategic leader, you will own integration strategy and execution across multiple domains, managing partner requirements, technical constraints, and external commitments while advancing platform modernization initiatives. You will partner closely with external stakeholders, understand their technical constraints and success criteria, and communicate risks and feasibility clearly in both technical and business terms. Responsibilities: </
About the Team The Plugin Developer Platform team builds the APIs, SDKs, and tools that let people extend ChatGPT and Codex. We work on plugins, connectors, the Model Context Protocol (MCP), and interactive apps. We want anyone to be able to turn a useful workflow into a plugin, share it, and have other people use it. A plugin can package instructions and skills with connections to the tools and data it needs. Our work covers plugin creation and publishing, the systems that run plugins across our products, and open standards that developers can build on. About the Role We’re looking for platform-minded engineers who know what it takes to build a platform developers want to use. You’ll work across developer-facing interfaces, APIs, and backend systems. You’ll own features from the first developer conversation through implementation and release. You’ll talk directly with developers, partners, and the open-source community. Their experience will inform the APIs and abstractions you design, the problems you prioritize, and the tradeoffs you make. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. What You’ll Do Design and ship APIs, SDKs, and services that developers use to extend ChatGPT and Codex. Make plugins easier to create, test, publish, update, and share. Improve compatibility and consistency across ChatGPT and Codex, including interactive app experiences. Contribute to MCP and other open standards, bringing practical developer needs into their design. Work with developers and partners to understand recurring problems and improve the platform, tooling, and documentation. Work with Product, Research, Security, and Trust & Safety on permissions, compatibility, and safe, reliable execution. You Might Thrive Here If You Have built software that other developers use. Your experience might include an open-source project, an API or SDK, a developer platform, internal too
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About Replit Replit is building the world’s most ubiquitous AI coding agent. Replit Agent can be used by anybody to bring their ideas to life. Whether it’s an app for yourself, the next great startup idea, or a tool to make you more productive at work, Replit Agent can help build it. Replit is also the leader in secure vibe coding. We protect apps, give users features to manage security risks, and help them vibe code more safely. About the Role In this role you will build powerful tools that help product engineers iterate rapidly on the Agent experience and directly enhance the core Agent itself. You’ll bridge the gap between the AI team (working on the core Agent logic) and the UX team (crafting delightful Agent experiences), enabling both groups to excel within their specialties. This role blends systems engineering, developer experience and product engineering. We tackle complex challenges across the full stack, from browser-based interfaces to high-performance backends to Linux systems engineering. We’re looking for engineers who have a keen sense of the product experience and how to power it with performant systems. On this team, you’ll have the opportunity to grow your skills across our infrastructure and product, and to lead end-to-end efforts with meaningful impact. We value diverse perspectives and encourage candidates from all backgrounds and experiences to apply. You Will Build high-throughput backend applications and services, like streaming chat between user and agent. Design a collaborative "Multiplayer Computer" that lets humans and AI agents work together on shared shells, filesystems, and state—conflict-free and in real time. Develop infrastructure (frontend & backend) that empowers product enginee
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the Team Product Platform builds and owns the shared foundations the rest of Replit is built on, spanning the full stack so every other team can ship features safely and quickly: backend infrastructure, connectors, product primitives, and the frontend platform. Our work is high-leverage and horizontal: when our foundations are solid every other team moves faster, and the role gives you exposure across the whole of engineering. We are a small, collaborative team that values curiosity and clear thinking over pedigree, and we work in the open by bringing each other the problem rather than just the request. We care more about how you reason and build than the route you took to get here. About The Role As a Product Engineer , you can focus on frontend, backend, or full-stack work building the shared systems other teams depend on. The work is guided by a few simple questions: Are our shared systems fast, reliable, and cost-efficient as traffic grows? Are we making product development safe by default, consistent, and faster? Can a builder connect a third-party service once and have it work safely across every app they build? Are user-facing surfaces consistent and fast, with shared primitives teams can build on? Is our codebase easy to navigate, change, and extend, including for AI coding agents? What you’ll do Design reusable primitives and interfaces with clear contracts and documentation that other teams adopt Work directly with product teams to turn their friction into platform improvements Profile and instrument shared systems, then ship the improvements that move latency, cost, and reliability Harden systems against failure and abuse, and make safe defaults the path of least resistance Set technical direction in a
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. About the role: We are seeking talented distributed systems engineers who are passionate about building innovative solutions for application deployment. Your mission will be to enhance the capabilities of Replit Infrastructure, optimize performance across global regions, and drive efficiency while delivering an exceptional user experience. If you have a strong foundation in software development, a deep understanding of cloud technologies, and a track record of delivering high-quality code, we want to hear from you. In this role you will: Expand Replit's cloud infrastructure offerings: Launch new cloud products to be used by Replit Agent to build complex apps. Collaborate with cross-functional teams to design and implement these features, empowering developers with a comprehensive suite of tools to build and deploy their applications efficiently. Enhance reliability and scalability: Identify bottlenecks, optimize critical paths, and implement robust monitoring and alerting systems. Work closely with the SRE team to ensure high availability and minimal downtime. Enable our customers to seamlessly scale their applications to meet the demands of their growing user base. Improve utilization of cloud infrastructure: Analyze our infrastructure costs and identify opportunities for optimization. Implement strategies to reduce cloud expenses without compromising performance or reliability. This could involve techniques such as resource provisioning, auto-scaling, cost-aware scheduling, and data lifecycle management. Your efforts will directly contribute to the financial efficiency of our cloud services. Required skills and experience: Distributed systems: Track record of working with platform-as-a-service, distributed storage, o
From $85K/yr
As Datadog’s in-house product experts, the Technical Escalation Engineering (TEE) team plays a critical role in driving our global success. We enable our customers, from the world’s most innovative startups to the largest enterprises. Through deep technical expertise, relentless problem-solving, and exceptional customer engagement, we educate, guide, and troubleshoot, delivering high-impact solutions that shape the customer experience. Whether through hands-on technical call, in-depth fact findings meeting, or complex investigations, we set the gold standard for technical excellence and customer advocacy. As part of our TEE team, you’ll tackle the most challenging technical problems, collaborate directly with Engineering and Product to refine and evolve our platform, and mentor teams worldwide, elevating the technical bar at every level. As part of the Technical Escalation Engineering (TEE) team, you’ll operate at the heart of Datadog’s ecosystem, working at the intersection of Technical Solutions, Engineering, Product, and our Customers. Every challenge you take on will directly impact the performance, scalability, and success of both our clients and our platform. You’ll be in an environment that moves fast, challenges you daily, and rewards curiosity, ownership, and technical excellence. This is your chance to shape the future of observability and security, driving innovation, mentoring teams, and influencing product direction while witnessing your expertise make an immediate and lasting impact. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Develop deep technical expertise and continuously learn as the product evolves. Investigate complex escalations, lead high-stakes technical calls, and drive solutions for our most critical customer cha
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. GitLab features must ship successfully across four distinct deployment platforms—Self-Managed, Dedicated, Dedicated for Government, and GitLab.com. Today, developers routinely discover platform compatibility issues late in development, triggering costly rework, delays, and emergency escalations. The Platform Readiness team is building an internal developer tooling capability that embeds platform readiness intelligence directly into GitLab's software development lifecycle to solve this problem. As an Intermediate Backend Engineer on the newly established Platform Readiness functional team in Developer Experience, you'll contribute
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software. * Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab. GitLab features must ship successfully across four distinct deployment platforms—Self-Managed, Dedicated, Dedicated for Government, and GitLab.com. Today, developers routinely discover platform compatibility issues late in development, triggering costly rework, delays, and emergency escalations. The Platform Readiness team is building an internal developer tooling capability that embeds platform readiness intelligence directly into GitLab's software development lifecycle to solve this problem. As a Senior Backend Engineer on the newly established Platform Readiness team, you'll architect and evolve the systems that enable feature
From $293.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Who We Are: Every day, tens of millions of people from around the world come to Roblox to play, learn, work, and socialize in immersive digital experiences created by the community. Our vision is to build a platform that enables shared experiences among billions of users. This is what’s known as the metaverse: a persistent space where anyone can do just about anything they can imagine, from anywhere in the world and on any device. Join us and you’ll usher in a new category of human interaction while solving exceptional challenges that you won’t find anywhere else. What You’ll Do: Robux is Roblox’s Virtual Currency: it’s how Users buy virtual items and how Creators earn a living. The Economy Platform team is the owner for all systems powering Robux transactions and has the goal of maintaining Roblox’s Economy healthy and vibrant. As a Principal Software Engineer for Virtual Economy Platform, you will be responsible for improving and scaling Roblox core monetization systems. You will report to Virtual Economy’s Senior Engineering Manager. The scope of Economy’s team includes the several Purchasing, Pricing and Payout systems. Those systems enable developers to monetize their game
From $278.5K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform team, you will design, scale, and maintain the foundational infrastructure powering our entire machine learning ecosystem. We are looking for accomplished engineers to spearhead the development of our next-generation ML tooling and platform capabilities. You will: Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator. Set technical strategy and oversee development of high scale and reliable infrastructure systems. Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle. Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources. Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices. Partner across organizations to build tooling, interfaces, and visualizati
About the Team OpenAI’s Platform team powers how millions of developers and enterprises build with our models. We provide APIs and agentic solutions used by global startups and fortune 500s. We work closely with product, engineering, design, and go-to-market to build a world-class platform that pushes the frontier of AI capabilities. About the Role As a Data Scientist on the Platform team, you will drive a data-driven culture for OpenAI’s API and B2B solutions. You’ll define the metrics that matter for developer success and enterprise value, measure the impact of new models and features, and partner with PMs and engineers to improve model quality, reliability, latency, and cost. Your work will shape how thousands of products adopt agentic AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will Embed with the Platform product team as a trusted partner, uncovering ways to improve developer experience, reliability, and usage growth Define north-star metrics across the developer funnel (activation, retention, growth), as well as latency/cost guardrails for new features and models Design and interpret A/B tests and controlled rollouts (e.g., new model versions, pricing/limits, new API features, new B2B products) Build source-of-truth dashboards and self-serve data tools for product, engineering, and go-to-market teams Translate product learnings into actionable feedback for Research (e.g., failure modes, eval gaps, model response quality) You might thrive in this role if you have 5+ years in a quantitative role in ambiguous, high-growth environments (platforms, APIs, or B2B products a plus) Depth in SQL and Python, with a track record proposing, designing, and running rigorous experiments Experience defining and operationalizing metrics from scratch (including reliability/latency/cost and safety) Strong cross-functional communication with PMs, enginee
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Senior Software Engineer — Cortex Training The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput. The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it. YOU WILL: Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane. Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional G
Other cities to consider
More places hiring for this role
Get new platform architecture manager jobs in United States by email
Daily job updates · Unsubscribe anytime