Jobs in United States

Aws And Tooling Platform Lead in United States

2,026 active opportunities · Updated October 2026

Explore current aws and tooling platform lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.

R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $278.5K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. ML Platform @ Roblox today supports hundreds of ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform team, you will design, scale, and maintain the foundational infrastructure powering our entire machine learning ecosystem. We are looking for accomplished engineers to spearhead the development of our next-generation ML tooling and platform capabilities. You will: Bootstrap and maintain Kubernetes and Cloud infrastructure for ML Platform components--Serving Layer, Metadata Store, Model Registry, and Pipeline Orchestrator. Set technical strategy and oversee development of high scale and reliable infrastructure systems. Propose and implement new platform tooling to improve time to production for MLEs and Data Scientists across the full ML lifecycle. Work on infrastructure projects such as GPU fleet management, hybrid-cloud orchestration, and writing custom Kubernetes controllers and resources. Stay abreast of industry trends in machine learning and infrastructure to ensure the adoption of leading-edge technologies and practices. Partner across organizations to build tooling, interfaces, and visualizati

AWSGCPDockerKubernetes
M
📍 United States· Full-time
✓ High-confidence listingCompany trend -93.7%

From $168K/yr

Quick readStrong listing-quality and freshness signals

MongoDB’s Developer Productivity organization exists to help engineers build and deliver high-quality software through a highly effective software development process and a strong foundation of shared tools and services. We are looking for a Senior Director to lead our Pipeline team. This role is tasked with bringing together the major systems and experiences that power software delivery at MongoDB. The team’s mission is to provide a reliable, scalable, secure, and effective platform for ensuring fast software deployability, leveraging AI native approaches. We are open to in-office, flexible or remote hiring across the US. The Team The Pipeline organization sits within Developer Productivity and is responsible for the systems, services, and user experiences that define MongoDB’s software delivery ecosystem. This is mission-critical infrastructure operating at substantial scale and supports a variety of software product delivery needs. Success in this role requires excellent product judgment for developer-facing experiences, strong systems and platform leadership, and the ability to align multiple teams around a cohesive strategy. Candidate Profile We’re looking for a senior engineering leader who can unify product-minded developer tooling with deep platform and operational excellence. The right candidate is passionate about developer productivity and has a track record of leading managers and teams through organizational growth, technical complexity, and cross-functional change. They should be comfortable owning a broad portfolio that spans developer experience, reliability and scale, release systems, telemetry, and operational health. They should also be able to work effectively with senior leaders and partners across engineering and product to set direction, allocate resources, and make trade-offs that balance near-term delivery with long-term platform function. The right candidate for this role will 12+ years of hands-on software engineering experience bui

MongoDBAWSAzureKubernetes
V
📍 United States· Full-time
✓ Quality checkedCompany trend -88.6%

At Vanta, our mission is to help businesses earn and prove trust. We believe that security should be monitored and verified continuously, and we empower companies to practice better security and prove it with ease. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. Vanta’s Developer Experience team builds the tools engineers use every day to bring ideas to production rapidly and reliably. You’ll empower other Vanta engineers to leverage cutting-edge technologies and best practices to make Vanta more performant and scalable on a platform level. Example projects include modernizing our CI/CD pipelines, introducing new test frameworks, launching AI-powered dev tools, and scaling developer environments to support a growing engineering team. This team has a wide breadth of impact across all of product engineering. The work we do compounds in value by making it easier for engineers to diagnose and solve bugs, streamline workflows, and ship value to our customers quickly and safely. Vanta engineers design and develop new product functionality and infrastructure leveraging modern frameworks and tooling, including TypeScript, React, Node.js, MongoDB, Github Actions, and various AWS services such as Fargate and ECS. Vanta has a kind and talented team, and while some have prior security experience, many have been successful at Vanta without it. We’d love for you to join us! You will: Set direction for critical dev infrastructure, enabling us to stay ahead of continued rapid growth Design and build CI and build systems that ensure Vanta engineers can develop and ship robust products quickly and confidently Improve the efficiency and reliability of our deployment workflows, including tools for hotfixes, rollbacks, and incident mitigation Lead development of tools that accelerate feedback loops — from typechecking and linting to running tests and deploying changes Build and maintain scalable developmen

TypeScriptReactNode.jsMongoDB
A
📍 United States
✓ High-confidence listingCompany trend +365.2%
Quick readStrong listing-quality and freshness signals

Abbott is a global healthcare leader that helps people live more fully at all stages of life. Our portfolio of life-changing technologies spans the spectrum of healthcare, with leading businesses and products in diagnostics, medical devices, nutritionals and branded generic medicines. Our 122,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: Position Overview The AI Platform Engineer builds and operates the machine learning and generative AI platform used by teams across Abbott Cancer Diagnostics. You'll own the full model lifecycle in production — data and feature pipelines, training and experimentation, evaluation and promotion, serving, and monitoring — along with the platform services, compute and tooling underneath it. This is hands-on infrastructure work backed by solid platform engineering practice: making inference fast and cheap, making the path from experiment to production repeatable and auditable, and shipping interfaces other engineers can build on — in support of software that ultimately reaches patients. Essential Duties Include, but are not limited to, the following: Build and maintain data, feature, and training pipelines for ML and LLM workloads — ingestion, transformation, fine-tuning, distributed training, and reproducible experiment execution with lineage tracked from dataset and code to resulting model. Implement automated evaluation and promotion gates — performance benchmarks, regression checks, and validation criteria that determine whether a model advances toward production. Automate the model lifecycle end to end through CI/CD and GitOps: packaging, promotion across environments, progressive rollout, and rollback. Build and operate production model-serving infrastructure for LLMs and predictive models, including inference optimization, autoscaling,

PythonJavaAWSKubernetes
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We believe Plaid has the power to be the next-gen Credit Bureau - supporting large scale adoption of cash flow into the credit underwriting process. The Credit Decisioning platform team is responsible for building best-in-class cashflow based insights products that enable lenders to make more holistic lending decisions and empower broader access to Credit products for prospective borrowers. We own the systems and tooling that form the platform to build and serve these insights at huge scale, partnering with our Data partners to release new products yearly. You will be defining the future architecture of Credit insights products and executing against an ambitious product roadmap. You will partner with our Product, Data Science, and Machine Learning team to iterate on and productionize new insights that enable our customers to make more holistic lending decisions. Responsibilities: Leading technical architecture and execution across credit insights products: everything from data fetching and online feature serving for API requests, to offline production pipelines and tooling for model training. Scaling and evolving the architecture through an expected ~100x increase in load from deterministic factors

AWSMachine LearningAI
S
📍 United States· Full-time
✓ Quality checkedCompany trend -81%

Synthesia is the world’s leading AI video platform for business, used by over 90% of the Fortune 100. Founded in 2017, the company is headquartered in London, with offices and teams across Europe and the US. As AI continues to shape the way we live and work, Synthesia develops products to enhance visual communication and enterprise skill development, helping people work better and stay at the center of successful organizations. Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion. Our total funding exceeds $530 million from premier investors including Accel, NVentures (Nvidia's VC arm), Kleiner Perkins, GV, and Evantic Capital, alongside the founders and operators of Stripe, Datadog, Miro, and Webflow. Remote (US East Coast preferred, for timezone coverage) About the team Cloud Infrastructure owns the platform every Synthesia product runs on — AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them. We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for. The role We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure. Today, too much critical operational knowledge — vendor relationships, cost management, and incident response — lives with one or two people. Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run. This is not simply a ticket-queue or keep-the-lights-on role. You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring . We deliberately pair operational and engineering work so the role grows rather than narrows. What you'll own Incident management & operational excellence — take custody of the incident process: on-call quality, resp

PythonMongoDBAWSKubernetes
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $345K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. We are seeking a visionary Principal Software Engineer to join our Compute organization and provide technical leadership for our Kubernetes infrastructure. You will drive the evolution of a platform that powers our global scale operations, transforming Kubernetes into a secure, reliable, and invisible foundation for our developers across our on-prem and public cloud fleet. Your mission is to balance cutting-edge innovation with rigorous platform stability, ensuring that internal and external customers have a seamless, high-performance experience at massive scale. You Will: Architectural Leadership: Serve as a technical lead for our Kubernetes ecosystem, setting the long-term architectural strategy for a platform that manages thousands of nodes and supports millions of concurrent requests. Customer Focus: Think deeply about how our internal customers consume and interact with compute, designing intuitive interfaces and tooling that simplify consumption of complex infrastructure services. Deep-Dive Engineering: Leverage deep expertise in Kubernetes internals, including custom controllers, operators, API server architecture, and etcd, to solve complex scaling bottlenecks and optimize our contr

AWSKubernetesGitMicroservices
C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

£225K – £325K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Manager of Security Engineering, your key responsibilities include: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Execute the long-term vision for the Security team in alignment with Cohere’s product and business goals. Collaborate closely with leadership to prioritize high-impact initiatives and strategic customer engagements. Vulnerability Management: Develop and implement enterprise-wide vulnerability management processes and tooling, including identification, prioritization, remediation tracking, and reporting, including customer artifacts Static Application Security Testing (SAST): Establish SAST programs, integrate tools into CI/CD pipelines, and analyze results to identify and remediate security flaws in source code Dynamic Application Security Testing (DAST): Implement DAST methodologies, configure scanning tools, and conduct regular assessments of running applications Penetration Testing: Lead and oversee internal and external penetration testing engagements, including web application, API, network and agentic AI platform inclu

PythonAWSAzureGCP
P
📍 New York, New York, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team. As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity. What excites you Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling. Architect and manage the SLO and error-budget

AWSKubernetesAIGo
P
📍 New York City, New York, United States· Full-time
✓ Quality checkedCompany trend -85.7%

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: The Experience team is at the center of one of the most exciting transitions in software development history — the shift from human-driven to agent-driven product experiences. We own Pinecone's API, clients, authentication, revenue, and observability systems, and right now that means redesigning all of it for a world where AI agents are first-class users alongside humans. This is a wide-scope role. You'll own things end-to-end — from backend architecture to API design to SDK and web surfaces. You’ll be working closely with product, design, and other engineering teams to identify user needs and build the right thing, at the right abstraction level, at the right time. Along the way, you will be building high-leverage platform capabilities that accelerate Pinecone’s product development and user growth systems. We're looking for an engineer who sees this moment for what it is: a rare opportunity to shape how developers and agents interact with a category-defining product. You're not waiting to see how the industry figures out MCP, agentic workflows, and AI-native interfaces — you're already experimenting, already forming opinions, already building. You know that speed and leverage matter more than labor, and you've internalized AI-assisted development not as a productivity trick but as a fundamentally different way of working. Responsibilities: Pioneer our agent experience. Shape how AI agents interact with Pinecone — designing interfaces, protocols (MCP), and tooling that make Pinecone the easiest and most capable platform f

JavaReactAWSAzure
P
📍 United States· Full-time
✓ Quality checkedCompany trend -85.7%

About Pinecone Pinecone is the knowledge infrastructure for AI at scale. Its leading vector database and knowledge engine, Pinecone Nexus, power accurate, performant AI applications for more than 9,000 customers and 800,000 developers worldwide. Pinecone's mission is to make AI knowledgeable. Pinecone is based in New York and raised $138M in funding from Andreessen Horowitz, ICONIQ, Menlo Ventures, and Wing Venture Capital. About the Team and Role: The Experience team is at the center of one of the most exciting transitions in software development history — the shift from human-driven to agent-driven product experiences. We own Pinecone's API, clients, authentication, revenue, and observability systems, and right now that means redesigning all of it for a world where AI agents are first-class users alongside humans. This is a wide-scope role. You'll own things end-to-end — from backend architecture to API design to SDK and web surfaces. You’ll be working closely with product, design, and other engineering teams to identify user needs and build the right thing, at the right abstraction level, at the right time. Along the way, you will be building high-leverage platform capabilities that accelerate Pinecone’s product development and user growth systems. We're looking for an engineer who sees this moment for what it is: a rare opportunity to shape how developers and agents interact with a category-defining product. You're not waiting to see how the industry figures out MCP, agentic workflows, and AI-native interfaces — you're already experimenting, already forming opinions, already building. You know that speed and leverage matter more than labor, and you've internalized AI-assisted development not as a productivity trick but as a fundamentally different way of working. Responsibilities: Pioneer our agent experience. Shape how AI agents interact with Pinecone — designing interfaces, protocols (MCP), and tooling that make Pinecone the easiest and most capable platform f

JavaReactAWSAzure
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $295.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Data Engineering team at Roblox plays a crucial role in enabling the company's success by developing and maintaining highly leveraged Core Data Sets, frameworks, and tooling to support the growing demand for analytics. As a Principal Data Engineer, you will work to define the data ontology for all of Roblox, establish best practices and standards for data operations and lifecycle management, design and build analytics tooling and frameworks, and influence event instrumentation. Additionally, this role is highly cross-functional, requiring close collaboration with Data Science, Experimentation, and Machine Learning teams to understand customer requirements and analytics applications, as well as with Data Infrastructure and Storage teams to develop integrated solutions. Join us and be a part of a dynamic team driving innovation and growth at Roblox. You Will: Partner with Data Science, Data Platform, Product, and Engineering to collect requirements to define the data ontology for all of Roblox Lead and mentor a growing team of Data Engineers to support Roblox's ever-evolving data needs Design, build, and maintain efficient and reliable batch and streaming data pipelines to mod

SQLAWSAzureGCP
R
📍 San Mateo, CA, United States· Full-time
✓ High-confidence listingCompany trend -100%

From $243.3K/yr

Quick readStrong listing-quality and freshness signals

Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Data Engineering team at Roblox plays a fundamental role in enabling the company's success by developing and maintaining highly leveraged Core Data Sets, frameworks, and tooling to support the growing demand for analytics. As a Senior Data Engineer, you will work to define the data ontology for all of Roblox, establish standard methodologies for data operations and lifecycle management, design and build analytics tooling and frameworks, and influence event instrumentation. Additionally, this role is highly multi-functional, requiring close collaboration with Data Science, Experimentation, and Machine Learning teams to understand customer requirements and analytics applications, as well as with Data Infrastructure and Storage teams to develop integrated solutions. Join us and be a part of a dynamic team driving innovation and growth at Roblox. This role will report to our Engineering Manager on the Data Engineering team. You Will: Partner with Data Science, Product, and Engineering to collect requirements to define the data ontology for all of Roblox Lead and mentor a growing team of Data Engineers to support Roblox's ever-evolving data needs Design, build, and maintain efficient an

SQLAWSAzureGCP
N
📍 Remote, United States· Remote
✓ High-confidence listingCompany trend -8%
Quick readStrong listing-quality and freshness signals

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions. We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly. What you'll be doing: Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support. Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry. Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data

PythonSQLAWSAzure
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Growth team drives user and revenue growth across ChatGPT’s consumer and business segments as well as other OpenAI products worldwide. We operate across the full funnel - from awareness and acquisition through activation, retention, and expansion - using a combination of global performance marketing, AI-powered workflows, in-product optimization, insights, experimentation, and creative ops engineering. About the Role We are hiring a Lifecycle Lead to build the company-wide owned-channel capability that helps teams reach users with relevant, timely, and trustworthy experiences. This senior, hands-on leader will set the lifecycle strategy, partner with Engineering to build the orchestration and deployment platform, and establish the operating model that allows teams across the company to launch and improve evergreen programs safely at scale. You will sit at the intersection of platform, product, and campaign strategy. You will partner with Engineering, Product, Data Science, and Analytics on the underlying systems, and with Product Marketing Managers and other client teams to design journeys that help new, active, and returning users reach value and build durable habits. In this role, you will: Partner with product to set the company-wide vision, roadmap, and operating model for lifecycle and owned-channel engagement. Partner with Engineering, Product, Data Science, and Analytics to shape the tooling and infrastructure for identity, audiences, eligibility, consent, triggers, orchestration, decisioning, frequency, experimentation, localization, quality assurance, and observability. Define scalable deployment workflows—including self-service and centrally supported paths, intake, templates, approvals, governance, service levels, and incident response—so teams across the company can launch safely and efficiently. Partner with Product Marketing Managers and other client teams to translate audience, product, and business goals into evergreen journey stra

AWSRestAIGo
🔔

Get new aws and tooling platform lead jobs in United States by email

Daily job updates · Unsubscribe anytime