About the Team The Ads Support Delivery team is responsible for helping successfully operate and grow on our Ads product. This includes technical guidance, troubleshooting complex delivery and monetization issues, and partnering closely with Product, Engineering, Trust & Safety and Go-To-Market teams to resolve customer-impacting problems and improve the platform over time. The team’s mission is to deliver a high-quality customer experience at scale by combining strong human support with automation, self-service, and AI-enabled workflows, while maintaining high operational rigor. About the Role: As a Support Delivery Lead for Ads, you will lead a team responsible for end-to-end support delivery across the ads ecosystem, including campaign setup, delivery, billing, measurement, and policy navigation. You will set the operational bar for quality, responsiveness, and consistency; coach and grow the team; and translate support signals into actionable improvements with Engineering, Product, and Go-To-Market partners. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead and support a team of Ads support engineers, ensuring they have the tools, clarity, and coaching needed to operate at a high bar in a technically complex domain. Set clear expectations and operating standards, run recurring performance reviews, and build development plans that grow both technical depth (ad tech fluency) and customer-facing excellence. Design and continuously improve support coverage for ad buyers, ensuring the team can diagnose delivery issues and monetization and integration issues with equal rigor. Act as the bridge between Support Delivery, Engineering, Product, and Go-To-Market teams. Drive alignment on priorities, escalation paths, launch readiness, tooling improvements and mechanisms to reduce repeated customer pain points. Partner with engineering teams
Jobs in United States
Aws And Tooling Platform Lead in United States
2,026 active opportunities · Updated October 2026
Showing
15 jobs
Explore current aws and tooling platform lead jobs across United States. Filter by work mode, employment type, experience, department, date posted and distance.
From $345K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Senior Engineering Manager, Safety Platform You will lead engineering pods within the Safety Platform organization, driving the technical vision and execution for the foundational systems that power every safety workflow at Roblox. You will own the core safety platform end-to-end: from the shared infrastructure and APIs that enable trust & safety capabilities across the company, to the tooling ecosystem that empowers internal operators to investigate, intervene, and resolve issues at scale. The scope also includes the Safety agentic platform — AI-powered systems that automate and augment safety workflows across detection, enforcement, and review pipelines. As the platform layer beneath every safety surface, your work will define how quickly and reliably Roblox can respond to emerging threats. This role requires a leader with a strong platform mindset who can balance technical rigor with broad organizational impact — working across User Safety, Trust & Safety Policy, Data Science, and Machine Learning to deliver scalable, extensible, and high-availability systems for one of the world's largest platforms. You Will: Lead and Develop: Recruit, hire, mentor, and inspire a diverse team of
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and detail-oriented GRC (Governance, Risk, and Compliance) Manager to build, support, and continuously enhance Baseten’s security governance, compliance, and privacy programs. As one of the early members of our security organization, you will play a key role in ensuring our platform meets and exceeds the highest standards for privacy, trust, and regulatory compliance. In this role, you’ll work cross-functionally with engineering, operations, legal, and leadership teams to develop policies, manage audits, and implement controls aligned with frameworks such as SOC 2, ISO 27001, ISO 27701, and FedRAMP. You’ll be instrumental in building scalable processes to manage risk, support customer assurance, and uphold Baseten’s commitment to security and compliance as we grow. RESPONSIBILITIES Governance & Policy Development: Design, implement, and maintain security governance frameworks, policies, and procedures that align with Baseten’s risk posture and industry best practices. Risk Management: Build and manage the company-wide risk assessment program, identifying, tracking, and mitigating key security and compliance risks. Compliance Operations: Lead efforts to achieve and maintain compliance with SOC 2, ISO 27001/27701, HIPAA, FedRAMP and other applicable standards and regulations. Audit & Certification Management: Coordinate external audits and certification processes, ensuring e
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead
From $293.8K/yr
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. Roblox’s database team develops the next-generation, multi-tenant database platform that elastically scales and underpins every online data workload at Roblox. As a principal engineer on the database team, you will shape the architecture, build and launch critical database capabilities that keep our services fast, reliable and efficient at global scale. You will report to the Technical Director for Storage. You will: Design and implement new engine features —indexing, storage formats, WAL and replication protocols, sharding, and query-planner enhancements—that push latency, throughput, and availability boundaries. Evolve the control plane to deliver elastic scaling, autonomous healing, and zero-downtime schema or tenant moves across global regions. Profile and optimize critical code paths using kernel-level tracing and advanced performance tooling; drive systematic tail-latency reductions. Establish engineering best practices by leading design reviews, performance benchmarks, failure drills, and post-incident retrospectives. Automate everything : develop frameworks for testing, CI/CD, rollout safety, observability, and autoscaling so that the platform operates hands-off at scale. Ment
From $182K/yr
Location Details: At GoDaddy the future of work looks different for each team. Some teams work in the office full-time, others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you’ll be working remotely from your home. You may occasionally visit a GoDaddy office to meet with your team for events or meetings. This position is not eligible to be performed in Alaska, Mississippi, North Dakota, or the Virgin Islands. GoDaddy is not currently considering candidates for this role in California, Seattle, or NYC. Join Our Team GoDaddy is hiring a Staff Software Engineer to help define and scale our Internal Developer Platform —a centralized system that powers how engineers across the company build, deploy, and operate software. This platform is used by 1,000+ engineers to manage everything from cloud infrastructure and application lifecycle to security, compliance, and cost transparency. In this role, you’ll operate as a technical leader at platform scale, shaping the architecture and direction of systems that directly impact engineering velocity across the company. You’ll work on high-impact initiatives such as AI-powered developer tooling, next-generation API platforms, and the evolution of a unified developer experience spanning APIs, CLI, and UI. This is a high-ownership, high-visibility role where Staff Engineers drive decisions, influence product direction, and partner across infrastructure, security, and platform teams. If you’re motivated by building systems that improve how other engineers work—and want to have a measurable impact on developer productivity at scale—this team sits at the center of GoDaddy’s engineering ecosystem. What You’ll Get to Do... Design and build platform services that power GoDaddy’s internal developer ecosystem, used by 1,000+ engineers. Lead architecture and technical direction for high-scale APIs, infrastructure orchestration,
Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: The Senior Platform Engineer II, AI Tooling on the Developer Experience Team (part of Foundation Engineering group) will lead the design and delivery of the internal AI platform that makes Drata's engineers more efficient - the tools, agents, and integrations that turn AI coding agents (and the rest of the AI dev stack) into a paved road for everyday engineering work. This is not a
About the role We’re looking for an engineering manager to lead a team building software systems that detect and prevent harmful misuse of frontier AI models—before incidents occur. This is a builder’s role: you’ll lead engineers shipping production services, detection pipelines, and mitigation mechanisms that protect frontier model integrity and reduce high-severity misuse risk. While this work intersects with frontier model development, security and risk, we’re explicitly seeking someone with a software engineering foundation who is comfortable building reliable systems that can operate at billions of users scale. In this role you will: Lead a team of software engineers building detection + mitigation systems for frontier model misuse, with an emphasis on model IP protection / distillation detection and emerging risk surfaces from autonomous agents. Set the technical roadmap and execution strategy: prioritize, design, ship, iterate, measure impact. Build production systems: services, pipelines, tooling, instrumentation, and automation that scale with frontier model usage. Partner deeply with Research and Product to translate evolving model capabilities into concrete tests, signals, and mitigations that can be deployed at scale. Drive strong engineering fundamentals: architecture, reliability, monitoring, performance, and operational excellence. Hire and grow an exceptional team across backend, data systems, and applied ML engineering domains as needed. Anticipate what breaks at scale as agentic workflows become more capable. You might thrive in this role if you: Experience building systems in adversarial, fast-evolving environments Are comfortable with ambiguity and novelty Have experience adjacent to security (e.g., abuse prevention, fraud, integrity, platform defense, auth/identity, malware/spam, adversarial environments) Communicate clearly and build trust quickly with senior stakeholders—pragmatic, collaborative, and calm under scrutiny. Significant experience
$342K – $445K/yr
About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The Role: SoFi's Cyber Defense organization is looking for an Offensive Security Lead to mature and grow our Penetration Testing and Red Team functions. This is a hands-on leadership role for someone who has spent years both doing the work and building the program around it — someone equally comfortable running a red team engagement against a critical banking platform and designing the operating model that lets a small team of offensive operators keep pace with a fast-growing fintech. A defining part of this role is modernizing how the team scales. We're looking for a leader who has already built and implemented AI-assisted penetration testing and red teaming programs — using AI tooling to accelerate reconnaissance, exploit development, attack-path analysis, and reporting — and who can bring that experience to bear on the program. You'll own the strategy, staffing, tooling, and execution quality of both disciplines, report into Cyber Defense leadership, and act as a trusted advisor to engineering, product, and risk partners across the company. What You’ll Do: Lead and unify Penetration Testing and Red Team into a single, cohesive Offensive Security function — shared standards, shared tradecraft, shared reporting, distinct missions. Set and execute the offensive security roadmap, aligning testing scope
About the Team The Startup Growth team is building the systems, programs, and automation that enable OpenAI to win startups at scale. We need to make it dramatically faster and easier to turn high-potential growth ideas into repeatable, measurable GTM motions without bespoke infrastructure for every program. About the Role The GTM Growth Programs Lead will own the shared capabilities that enable Startup GTM to incubate, launch, measure, and scale growth programs. The role will initially own credit and commercial programs while using that work to build reusable infrastructure for a broader portfolio of GTM motions. The goal is not to centralize program ownership: teams closest to startups should continue to identify, incubate, and own growth motions. This role will build the operating system that makes them dramatically better and faster at doing so. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build the GTM program platform for Startups. Develop shared infrastructure, tooling, and agentic systems that turn growth opportunities into scalable programs, automate work across the program lifecycle, and make each subsequent motion faster to launch. Own credit and commercial programs. Drive credit/commercial strategy and evolution while turning capabilities such as audience selection, eligibility, offer configuration, distribution, and measurement into reusable building blocks. Create a repeatable path from idea to scaled motion. Turn “we need to increase X” into targeted, operationalized, measurable GTM motions. Build leverage throughout the team. Enable ADs, VC Partnerships, and other Startup GTM teams to incubate and own programs independently rather than routing every motion through centralized operations. Develop shared measurement and experimentation frameworks and capabilities. Establish common approaches to sizing, economics, attribution
From $265K/yr
We're transforming the grocery industry At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it together. Where others see a simple need for grocery delivery, we see exciting complexity and endless opportunity to serve the varied needs of our community. We work to deliver an essential service that customers rely on to get their groceries and household goods, while also offering safe and flexible earnings opportunities to Instacart Personal Shoppers. Instacart has become a lifeline for millions of people, and we’re building the team to help push our shopping cart forward. If you’re ready to do the best work of your life, come join our table. Instacart is a Flex First team There’s no one-size fits all approach to how we do our best work. Our employees have the flexibility to choose where they do their best work—whether it’s from home, an office, or your favorite coffee shop—while staying connected and building community through regular in-person events. Learn more about our flexible approach to where we work. Overview Enterprise Foundations is the engineering team responsible for the architectural and integration work that allows Instacart's Retailer Platform to scale across the world's largest grocers. Our work spans omnichannel integrations — bringing products like FoodStorm catering and Caper Carts onto unified Instacart platform rails — enterprise extensibility (sandbox environments, configuration systems, retailer-facing tooling), and the cross-cutting data isolation work that allows Instacart's ML and Data Engineering teams to safely build per-retailer models on top of the platform. As Instacart's enterprise retail business expands to hundreds of partner brands and new international markets, this team's work is foundational to how the platform matures. We are seeking a Senior Engineering Manager to lead this team of ~12 engineers. You'll set the unifying technical
From $100K/yr
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is building high-performance AI systems and the supply chain required to deliver them at scale. We are looking for a Global Supply Chain Manager to own OSAT partnerships, backend manufacturing execution, and supply continuity from wafer-out through final shipment. This role is based in Santa Clara, CA; Austin, TX, or Toronto, ON, with regular travel to Taiwan and other Asia-based partner sites, approximately 25-30%. Candidates should be located near one of these hubs and able to work effectively across North American and Asia-Pacific time zones. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are Bring 7+ years of experience in semiconductor supply chain, manufacturing, or supplier management, along with a bachelor’s degree in Supply Chain, Engineering, or a related discipline. Have a strong understanding of OSAT workflows, including flip-chip, wire bond, and wafer-level packaging, with hands-on experience working with Taiwan-based OSAT partners. Are comfortable negotiating complex supplier agreements, pricing, tooling, or NRE costs and managing relationships across technical, commercial, and operational issues. Are a clear, direct, a
Other cities to consider
More places hiring for this role
Get new aws and tooling platform lead jobs in United States by email
Daily job updates · Unsubscribe anytime