Jobs in United States

Platform Product Partnerships Lead in United States

3,706 active opportunities · Updated October 2026

Explore current platform product partnerships lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

A
📍 United States· Full-time· Remote
✓ Quality checkedCompany trend -99%

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way. The Community You Will Join: BizTech fosters culture and connection at Airbnb by providing reliable corporate tools, innovative products, and technical support for all teams. We drive technical breakthroughs and strategies that redefine what it means to belong anywhere, delivering greater value for the business and our people. The Global Operations team at BizTech manages production services across Airbnb’s corporate environment, delivering reliable operations through Observability, Incident Management, Core Operations, and AI-enabled automation. We partner across BizTech to scale service quality, efficiency, and resilience. The Difference You Will Make: As a Senior Staff Engineer in Operations, you will lead and mentor a high-performing team to scale our AI-enabled operations model and deliver AIOps solutions that streamline operational workstreams and help BizTech teams focus on their core work with confidence. Ops owns triage and resolution, proactive monitoring across networks, systems, applications, and cloud services via a homegrown observability platform, and drives process excellence through automation and shift-left programs. You will set the technical bar, model operational excellence, and ensure high-quality, reliable service. Your scope includes leading projects ac

PythonAWSCI/CDAI
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a

PythonAWSRestMachine Learning
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training Performance Engineer, you’ll drive efficiency improvements across our distributed training stack. You’ll analyze large-scale training runs, identify utilization gaps, and design optimizations that push the boundaries of throughput and uptime. This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale. You’ll help ensure that our clusters are running at peak performance, enabling OpenAI to train larger, more capable models with the same compute budget. This role is based in San Francisco, CA. We use a hybrid work model of three days in the office per week and offer relocation assistance to new employees. In this role, you will: Profil

PythonAWSRestAI
N
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

$230K – $260K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About The Role Millions of people rely on Notion to do their most important work, and protecting that trust is foundational to everything we build. We’re looking for a hands-on Detection Engineer to build and operate the systems and workflows we use to detect and respond to attacks across Notion’s cloud-native environment. You’ll ship high-signal detections, improve the platform that powers them, participate in incident response, and help shape how detection and response engineering scales at Notion. You’ll work closely with Engineering, Corporate Security, and Infrastructure, with broad latitude to identify gaps, prioritize investments, and build what’s needed next. We view detection and response as a software engineering discipline: detections are code, platforms are products, and measurement matters What You'll Achieve Design and maintain high-signal detections across cloud, identity, endpoints, and SaaS environments. Build and improve the detection platform, including rule lifecycle management, tuning, measurement, and rollout safety. Develop tooling and automation that accelerate triage, enrichment, investigation, and detection

AWSAzureGCPKubernetes
N
📍 New York, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.6%

From $299K/yr

Quick readStrong listing-quality and freshness signals

Who We Are Notion is the collaborative AI workspace where teams and agents think together . We're building one place where your knowledge, projects, meetings, and AI tools live side by side, so work is faster, clearer, and less fragmented. Millions of individuals, small teams, and large companies run their work on Notion. Notinos (our employees) are customer zero in bringing this future of work to life. We care about craft, building things that last, and the belief that great work is still fundamentally human. Our goal isn’t to ship the next feature. Each and every team of Notinos is working to set the standard for how humans work together in the AI era. From building a business’s system of record to making and managing AI agents to automating away the busy work, we care deeply about giving our customers more time for their life’s work. About the Role We’re rolling out Go support at production scale at Notion, and we need an owner who can make it durable. You’ll lead the work to turn Go into a fully supported, well-operated platform: reliable and scalable service patterns, paved paths for our tooling stack, and the guardrails that make building in Go feel fast and safe. This role matters because our next wave of AI and agent-driven products will require backend services where Node won’t always be the right fit, and the platform decisions we make now will compound for years. While Go is the core focus, this role sits within Developer Experience and will regularly tackle other high-leverage engineering productivity challenges, developer experience ranging from AI-assisted development workflows and remote agent environments to CI performance, deployments, and reliability tooling. This role can be based in either San Francisco or New York City. We work from our offices on Mondays, Tuesdays and Thursdays (our Anchor Days) because we do our best thinking and building together in person. We’re looking for someone who’s excited to work alongside the team during those days.

TypeScriptKubernetesRestAI
S
📍 United States· Full-time
✓ Quality checkedCompany trend -93.3%

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. The Red Team in the Security Foundations org is responsible for conducting security assessments against Snowflake’s diverse cloud environment and finding vulnerabilities in software, systems, and networks. In this role, you will execute security assessments across Snowflake’s corporate and product environments. You will partner with cross-functional teams to communicate gaps discovered during engagements and help drive remediation. You will develop tools, perform security research, and build infrastructure to augment the Red Team’s capabilities. Our ideal candidate wakes up each morning thinking about different parts of the business that can benefit from security assessments. Their goal is to identify relevant security risks and help the business understand them so they can build effective defenses and protect Snowflake customers and their data. RESPONSIBILITIES Develop tools, methodologies, and infrastructure to support Red Team engagements in a variety of cloud environments and novel platforms Participate in security assessments against a diverse cloud environment and find vulnerabilities in software, systems, and networks Set scope, objectives, and timelines for security assessments and leverage data to create meaningful metrics Work with security and engineering teams t

JavaScriptPythonJavaAWS
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -73.5%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Team Description: Plaid is evolving into an AI-first company, and Intelligent Tooling sits at the center of that transformation. The Intelligent Tooling team is being built from the ground up, and our mission is to establish the technical foundations, operating model, and internal platforms that embed AI deeply into Plaid’s coding tools, internal systems, and the entire software development lifecycle. When we are successful, engineers across Plaid will delegate lower-leverage work to AI agents, move faster with confidence, and spend more of their time designing and inventing for customers. Intelligent Tooling owns the platforms and systems that make this possible - from AI coding integrations and SDLC agents to the internal tools that power Plaid’s operations. Role Description: As a Staff Software Engineer on the Intelligent Tooling team, you will build and operate internal systems that directly impact how engineers across Plaid do their work, and own the technical direction for major parts of that surface. This is a hands-on role with significant ownership, where success is measured by real adoption, reliability, and improvements to developer experience. You will work on AI-powered tooling, interna

AWSCI/CDRestAI
O
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -84.1%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal and VMs dependable & scalable foundations for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Software Engineer, Host Assurance to build and operate the services, APIs, and host software that establish and maintain trust in our compute infrastructure. You will own production software from design and implementation through testing, rollout, observability, and operation. Your work will support capabilities such as machine identity, certificate issuance and enrollment, secure bootstrap, and host attestation across bare-metal and VM environments. Success in this role requires strong technical judgment, the ability to reason across software and host-system boundaries and learn unfamiliar parts of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—inc

AWSRestAIGo
N
📍 Santa Clara, United States
✓ Quality checkedCompany trend -13.7%

NVIDIA is leading groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU -- our invention -- serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables groundbreaking creativity and discovery, and powers inventions that were once considered science fiction, including artificial intelligence to autonomous cars. We are the GPU Communications Libraries and Networking team at NVIDIA. We build communication libraries like NCCL, NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center platforms and scalable communications software. DL and HPC applications have a huge compute demands and already run at scales of up to tens of thousands of GPUs. GPUs are connected with high-speed interconnects (e.g. NVLink, PCIe) within a node and with high-speed networking (e.g. InfiniBand, Ethernet) across nodes. Efficient and fast communication between GPUs directly impacts end-to-end application performance. This impact continues to grow with the increasing scale of next generation systems. This is an outstanding opportunity to advance the state-of-the-art, break performance barriers, and deliver platforms the world has never seen before. Are you ready to build the new and innovative technologies that will help realize NVIDIA's vision? What you will be doing: Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems. Design and implement new communication technologies to accelerate AI and HPC workloads. Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects. Build proofs-of-concept, conduct experiments,

LinuxArtificial IntelligenceAI
D
📍 Massachusetts, New York, United States· Full-time
✓ High-confidence listingCompany trend -88.9%

From $192K/yr

Quick readStrong listing-quality and freshness signals

Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe. The platforms are optimized for durability, high availability, low latency, internet-scale footprint and operability. At Datadog, we place value in our office culture - the relationships that it builds, the creativity it brings to the table, and the collaboration of being together. We operate as a hybrid workplace to ensure our employees can create a work-life harmony that best fits them. What You’ll Do: Build fault-tolerant, horizontally scalable solutions running in multi-tenant environments Write in Go, Java Rust or C++, amongst other languages Use Kafka, Redis, Cassandra, Elasticsearch and other open-source components Own meaningful parts of our service, have an impact, grow with the company Who You Are: 6+ years of experience You have a BS/MS/PhD in a scientific field or equivalent experience You have significant backend programming experience in one or more languages (Go, Java, Rust, C++) You have been exposed to working on problems (high durability / low latency /…) You can get down to the low-level when needed You care about simple designs and performance You want to work in a fast, high-growth startup environment that respects its engineers and customers You have demonstrated ability to use AI coding tools in day-to-day workflows and validate, critique, and refine AI-generated output. Bonus: you’re motivated to push the boundaries of how AI can improve software engineering best practices and contribute to building AI-enabled products. This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. Datadog values peo

JavaRedisAIC++
O
Relocation support. Relocation assistance is stated. This does not establish visa sponsorship.
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -84.1%

About the Team Security is foundational to OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security organization protects OpenAI’s technology, people, and products by building and operating deeply technical systems that must work reliably at massive scale. Our work underpins OpenAI’s commitments around safety, privacy, and security across research, products, and emerging platforms. The Host Assurance team exists to make bare metal a dependable, scalable foundation for OpenAI: secure by default, verifiable in practice, and resilient across providers and operating models. We operate at the trust boundary between physical hardware and cloud-scale orchestration, ensuring that hosts are eligible to safely run workloads with predictable security properties and auditability. About the Role OpenAI is seeking a Security Engineer, Host Assurance to help build the trust foundations for bare-metal platforms across OpenAI’s global infrastructure. This is a deeply hands-on engineering role for a builder who can design, implement, and operate the core security infrastructure that establishes trust in hardware platforms before they are eligible to run workloads. Success in this role requires strong technical judgment, the ability to work comfortably at low levels of the stack, and a practical mindset for building systems that are secure, reliable, and usable in fast-moving production environments. The systems you build will sit on the critical path of OpenAI’s frontier infrastructure investments and will directly shape how large amounts of compute are brought online - securely, responsibly, and at global scale - underpinning long-lived commitments around privacy, security, and reliability. You will partner closely with infrastructure, research, and confidential computing initiatives—including novel hardware platforms and emerging deployment models– to make the secure path the easiest path. This role is well suited for engineers who enjo

AWSRestAgileAI
P
📍 New York, NY, United States· Internship
✓ High-confidence listing
Quick readStrong listing-quality and freshness signals

A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more. The Role Forward Deployed Infrastructure Engineers (FDIEs) build, operate, and maintain the infrastructure that powers Palantir’s platforms and production deployments. As an FDIE intern, you’ll work alongside full-time FDIEs to deploy and operate Palantir software across real production environments, automate manual processes, and develop novel solutions to infrastructure challenges using tools like Foundry and Apollo. Every day looks different — you might be debugging a distributed systems issue, building automation to replace a manual runbook, or designing infrastructure improvements that scale across multiple deployments. You’ll be treated as a full member of the team, with real ownership over the work you take on. Core Responsibilities As an FDIE intern, your responsibilities look similar to those at a small startup, with the resources, stability, and mentorship of an established tech company. You’ll work in small teams with minimal supervision and own end-to-end execution of real infrastructure projects. Your day might span discussing systems architecture with fellow engineers, debugging a production issue, building automation to eliminate a manual process, or deploying new Palantir products across production environments. FDIE interns are treated just like full-time engineers, with significant freedom and ownership over their work. Specifically, you can expect to: Deploy and operate Palantir software across production environments, including monitoring, alerting, configuration management, and upgrades Debug, improve, and optimize Palantir’s services and infra

JavaScriptTypeScriptPythonJava
D
📍 New York, New York, United States
✓ Quality checkedCompany trend -88.9%

Datadog's Software Engineers with Systems depth leverage their experience with systems and tooling to build software that ensures Datadog remains reliable, performant, and secure. For this track, their Software Engineering experience may resemble the Distributed Systems track, but is typically applied in combination with their systems experience to build and run internal platforms and tools that our products are built on. These people typically have deep experience building and managing large cloud infrastructure deployments, or leading reliability efforts for orgs similar to ours, or building release machinery to allow hundreds or thousands of devs to do their jobs without stepping on each others' toes. The systems and tooling where they may have experience depth may include (but not limited to): bazel, build tooling, cassandra, CDN, chef, configuration management, container orchestration, consul, docker, elasticsearch envoy, haproxy, kafka, kubernetes, load balancing, network architecture, postgres, redis, release management, RPC frameworks, service discovery, spinnaker, terraform, zookeeper. Bonus: You’re excited about leveraging AI tools to enhance how you code, solve problems, and build – or eager to learn how This job is available in various departments within our company; to conform to US export control regulations, some of these roles may require candidates to be eligible for any required authorizations from the US government. #LI-KM5 Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan. Th

PostgreSQLRedisDockerKubernetes
DC
📍 New York, New York, United States· Full-time
✓ High-confidence listing

From $131K/yr

Quick readStrong listing-quality and freshness signals

Role Overview You’re a seasoned Site Reliability Engineer who loves owning complex infrastructure, making things run faster, safer, and with less manual effort. In this Staff‑level role, you’ll design and operate VMware‑based private cloud platforms that power mission‑critical SaaS products used by customers around the world. You’ll work across Linux, Windows Server, networking, storage, and automation frameworks to increase reliability, reduce toil, and modernize a global datacenter environment. You’ll have the scope to set technical direction, build automation at scale, and mentor engineers while staying hands‑on with VMware vSphere, F5/AVI load balancers, and hybrid Active Directory. Here’s a breakdown of what you’ll do (not all of it, just the important stuff) Lead the architecture, deployment, and ongoing optimization of VMware vSphere–based private cloud infrastructure across multiple global datacenters. Design and build automation using PowerShell/PowerCLI, Ansible, Python, and CI/CD tools to streamline provisioning, configuration, and compliance. Administer, harden, and troubleshoot Linux (RHEL/CentOS/Ubuntu) and Windows Server environments that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and service accounts across hybrid on‑prem and cloud environments. Partner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. Document architectures and runbooks, participate in on‑call and change management, and mentor engineers while influencing long‑term reliability and automation strategy. These are the essentials you’ll need to get an interview 10+ years of experience in systems or infrastructure engineering, including operating large‑scale enterprise or SaaS datacenter environments. Deep hands‑on expertise with VMware vSphere (ESXi, vCenter, DRS, HA, vMotion, distributed switches) in production

PythonAWSAzureCI/CD
P
📍 San Francisco, California, United States· Full-time· Remote
✓ Quality checkedCompany trend -73.5%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam. The Integrations Operations Engineering (IOE) team strengthens Plaid's network with financial institutions by directly resolving integration issues to increase the reliability and quality of data, and by building new integrations to grow our financial network. We sit at the nexus of engineering, a deep understanding of Plaid's products and customers, and direct financial-institution relationships: we write and ship the code that keeps connectivity healthy, we use data to focus on the issues with the greatest customer impact, and we work directly with data partners (the financial institutions and platforms themselves) to resolve the problems that can't be fixed from our side alone. The team plays a mission-critical role in ensuring industry-leading connectivity so our customers can meet ever-expanding financial-services use cases and reach as many users as possible. What You'll Do Investigate and resolve the highest-impact integration issues by writing maintainable, tested code and deploying it to production, then monitor for regression or degradation after your changes ship. Prioritize by customer impact. The team runs a business-value-based prioritization model that automatically

TypeScriptNode.jsSQLAWS
🔔

Get new platform product partnerships lead jobs in United States by email

Daily job updates · Unsubscribe anytime