Jobs in United States

Aws And Tooling Platform Lead in San Francisco

866 active opportunities · Updated October 2026

Explore current aws and tooling platform lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

D
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -93.7%
Quick readStrong listing-quality and freshness signals

Drata is building the trust layer between great companies - automating compliance, managing risk, and helping organizations prove trust continuously as they scale. We're Dratanauts: a global crew of 600+ professionals united by a culture that rewards integrity, ownership, and raising the bar, no matter where in the world we're working from. Why Join the Drata Team? At Drata, you're not maintaining legacy compliance software - you're building the agentic AI platform defining what trust looks like for the next generation of companies. Here's what makes the work itself worth showing up for: Problems without a playbook: You'll work at the edge of AI and security, building agentic governance, continuous compliance, and real-time trust verification to solve problems that don't have an established answer yet. You're writing it as you go. Real ownership, not just process: Our values center on owning outcomes and raising the bar, not checking boxes. You're expected to have opinions and back them. A seat at the table: Your perspective is unique and valued. Open debate and diverse viewpoints are built into how decisions actually get made here, at every level. Growth at rocketship speed: Drata is scaling fast, which means scope grows fast too. High performers get more ownership, visibility, and experience. A crew, not just coworkers: Dratanauts consistently describe a "come as you are" culture with sharp, curious people—the kind of team that makes hard problems genuinely fun to solve. See what they say here and follow us on LinkedIn for company news, employee stories, and career updates. Job Summary: The Senior Platform Engineer II, AI Tooling on the Developer Experience Team (part of Foundation Engineering group) will lead the design and delivery of the internal AI platform that makes Drata's engineers more efficient - the tools, agents, and integrations that turn AI coding agents (and the rest of the AI dev stack) into a paved road for everyday engineering work. This is not a

TypeScriptPythonNode.jsSQL
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the role We’re looking for an engineering manager to lead a team building software systems that detect and prevent harmful misuse of frontier AI models—before incidents occur. This is a builder’s role: you’ll lead engineers shipping production services, detection pipelines, and mitigation mechanisms that protect frontier model integrity and reduce high-severity misuse risk. While this work intersects with frontier model development, security and risk, we’re explicitly seeking someone with a software engineering foundation who is comfortable building reliable systems that can operate at billions of users scale. In this role you will: Lead a team of software engineers building detection + mitigation systems for frontier model misuse, with an emphasis on model IP protection / distillation detection and emerging risk surfaces from autonomous agents. Set the technical roadmap and execution strategy: prioritize, design, ship, iterate, measure impact. Build production systems: services, pipelines, tooling, instrumentation, and automation that scale with frontier model usage. Partner deeply with Research and Product to translate evolving model capabilities into concrete tests, signals, and mitigations that can be deployed at scale. Drive strong engineering fundamentals: architecture, reliability, monitoring, performance, and operational excellence. Hire and grow an exceptional team across backend, data systems, and applied ML engineering domains as needed. Anticipate what breaks at scale as agentic workflows become more capable. You might thrive in this role if you: Experience building systems in adversarial, fast-evolving environments Are comfortable with ambiguity and novelty Have experience adjacent to security (e.g., abuse prevention, fraud, integrity, platform defense, auth/identity, malware/spam, adversarial environments) Communicate clearly and build trust quickly with senior stakeholders—pragmatic, collaborative, and calm under scrutiny. Significant experience

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%

$342K – $445K/yr

Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are seeking a Technical Lead to lead deployment and operations for OpenAI’s Silicon & Systems team. This person will become the Directly-Responsible Individual responsible for bringing OpenAI’s custom silicon and associated systems into data center environments, ensuring successful deployment, bring-up, validation, operational readiness, and ongoing reliability at scale. This role sits at the intersection of silicon, systems, infrastructure, data center operations, and software. You will lead a team focused on taking new hardware platforms from lab validation into production data center deployment. You will be responsible for building the operational processes, technical workflows, tooling, and cross-functional alignment required to deploy and operate custom AI hardware reliably in OpenAI’s supercomputing infrastructure. The ideal candidate is both a strong leader and a deeply technical operator. You should be comfortable staying close to the technical details of hardware bring-up, fleet deployment, debugging, system validation, data center integration, and production operations. This role requires strong execution, excellent cross-functional judgment, and the ability to drive clarity in ambiguous, fast-moving environments. In this role, you will: Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments Own the path from hardware bring-up and validation through production deployment, operati

AWSRestAIGo
S
📍 San Francisco, CA, United States· Full-time
✓ High-confidence listingCompany trend -90.5%
Quick readStrong listing-quality and freshness signals

Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. The Role: SoFi's Cyber Defense organization is looking for an Offensive Security Lead to mature and grow our Penetration Testing and Red Team functions. This is a hands-on leadership role for someone who has spent years both doing the work and building the program around it — someone equally comfortable running a red team engagement against a critical banking platform and designing the operating model that lets a small team of offensive operators keep pace with a fast-growing fintech. A defining part of this role is modernizing how the team scales. We're looking for a leader who has already built and implemented AI-assisted penetration testing and red teaming programs — using AI tooling to accelerate reconnaissance, exploit development, attack-path analysis, and reporting — and who can bring that experience to bear on the program. You'll own the strategy, staffing, tooling, and execution quality of both disciplines, report into Cyber Defense leadership, and act as a trusted advisor to engineering, product, and risk partners across the company. What You’ll Do: Lead and unify Penetration Testing and Red Team into a single, cohesive Offensive Security function — shared standards, shared tradecraft, shared reporting, distinct missions. Set and execute the offensive security roadmap, aligning testing scope

AWSAzureGCPRest
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Startup Growth team is building the systems, programs, and automation that enable OpenAI to win startups at scale. We need to make it dramatically faster and easier to turn high-potential growth ideas into repeatable, measurable GTM motions without bespoke infrastructure for every program. About the Role The GTM Growth Programs Lead will own the shared capabilities that enable Startup GTM to incubate, launch, measure, and scale growth programs. The role will initially own credit and commercial programs while using that work to build reusable infrastructure for a broader portfolio of GTM motions. The goal is not to centralize program ownership: teams closest to startups should continue to identify, incubate, and own growth motions. This role will build the operating system that makes them dramatically better and faster at doing so. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Build the GTM program platform for Startups. Develop shared infrastructure, tooling, and agentic systems that turn growth opportunities into scalable programs, automate work across the program lifecycle, and make each subsequent motion faster to launch. Own credit and commercial programs. Drive credit/commercial strategy and evolution while turning capabilities such as audience selection, eligibility, offer configuration, distribution, and measurement into reusable building blocks. Create a repeatable path from idea to scaled motion. Turn “we need to increase X” into targeted, operationalized, measurable GTM motions. Build leverage throughout the team. Enable ADs, VC Partnerships, and other Startup GTM teams to incubate and own programs independently rather than routing every motion through centralized operations. Develop shared measurement and experimentation frameworks and capabilities. Establish common approaches to sizing, economics, attribution

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a product manufacturing & quality engineer, who will be responsible for driving technical initiatives related to the manufacturing, quality and reliability of our AI supercomputer hardware systems to ensure product success from concept to launch and through mass production. You’ll have the opportunity to coordinate with functional SMEs and work with a wide range of stakeholders, from design engineering and operations teams, TPMs, external industry vendors and partners to ensure that all products are developed and delivered on time and to the highest quality standards. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees In this role, you will: Own the integrated manufacturing and quality readiness for a product across L6, L10, and L11, with clear gates, milestones, deliverables, owners, and closure criteria. Lead readiness of process flows, tooling, fixtures, assembly operations, test interfaces, and production controls. Review and contribute to work instructions. Translate product requirements into qualification plans, process controls, test requirements and acceptance criteria with design engineering and Area SMEs Coordinate and drive execution of product and process qualification, reliability testing, and validation with the relevant SMEs. Maintain traceable evidence that assigned products and processes meet agreed performance, reliability,

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team OpenAI’s Forward Deployed Engineering team partners with leading semiconductor companies to deploy production-grade AI systems across the entire chip design lifecycle: design, verification, and physical design. We operate at the intersection of customer delivery and core platform development, embedding deeply with customers to translate frontier model capabilities into systems that materially improve engineering workflows and accelerate innovation. Our work turns early, high-touch deployments into repeatable solution patterns, reference architectures, and evaluation practices that scale across the semiconductor ecosystem. About the Role We are seeking a highly skilled Physical Design Engineer to join our semiconductor-focused Forward Deployed Engineering team. This is a senior IC role that will begin with a strong emphasis on physical design expertise, technical judgment, advisory leverage, and customer credibility, with the expectation that the person will grow into a broader Forward Deployed Engineering role over time. In the near term, you will serve as the team’s physical design SME across semiconductor deployments: helping FDEs, Product, and Research understand backend implementation workflows, pressure-test AI-assisted solution ideas against real physical design constraints, and raise the quality of our customer-facing technical work. You will help the broader team build fluency in implementation flows, EDA tooling, signoff methodology, and the trade-offs that shape physical design decisions in practice. Over time, we expect this role to expand beyond SME support into broader FDE ownership: partnering directly with customers, shaping deployment strategy, building and iterating production-grade AI systems, driving technical workstreams, and helping turn high-touch semiconductor deployments into repeatable solutions. This is a strong fit for someone who brings deep physical design expertise today and is excited to grow into a customer-facing, syst

PythonAWSRestAI
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Description of the team The Dashboard Foundations team is the product platform team that stewards the Plaid Dashboard ( dashboard.plaid.com ). As both the portal for applying for production access and also the host for many products, the Dashboard is a critical touchpoint for our customers. Our mission is to build the platform of every product engineer’s dreams, with rich tooling, abstractions, and resources available to support every phase of the software development lifecycle so that developing a high quality, secure product is fast and easy. Our customers are over a dozen teams building in the Dashboard, representing products in areas such as Fraud, Credit, Signal, Account Verification, Transfer products, and more. We are a full stack team made up of former product engineers, drawing upon our experience to set our north star vision. We have in-person members in New York and San Francisco as well as some members distributed in various locations across the United States. Responsibilities You will lead a team of 8 engineers, ranging from Junior to Staff, developing them through clear goal setting, coaching, and feedback. You’ll define and drive the long-term strategy for this foundational area, in c

O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking Manufacturing Engineers to lead the development of processes, tooling, and prototype builds for custom motors and actuators. You will own a primary area as a Stator Process / Manufacturing Engineer, Tooling / Fixture Engineer, or Prototype Manufacturing Engineer, taking that work from early development through validation and repeatable execution in close partnership with mechanical, electromagnetic, electrical, test, quality, and supplier teams. These roles focus on the development, integration, and validation of electromechanical manufacturing capabilities, including stator winding processes, assembly and inspection tooling, and actuator prototype builds. You will help translate engineering designs into reliable hardware while establishing scalable processes, equipment, and build practices for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 5 days a week. In this role, you will Each opening focuses on one of the three specialties below, with shared responsibility for reliable processes and hardware. Stator Process / Manufacturing: develop and validate stacking, winding, termination, and potting processes. Establish process parameters, work instructions, and defect controls that produce consistent results across trained operators. Tooling / Fixture: design and deliver winding tools, assembly fixtures, inspection gauges, and bench equipment, from CAD and drawings through fabrication, commissioning, and troubleshooting. Improve setup time, labor, and repeatability. Prototype M

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role We are seeking a Mechanical Design Engineer to lead the development of motor mechanical designs and prototype hardware for advanced robotic systems. You will own stator and rotor mechanical development from early concepts and detailed design through hands-on builds and prototype validation, partnering closely with electromagnetic, electrical, test, and manufacturing teams. This role focuses on the design, integration, and validation of motor components and prototype processes, including laminations, stack assemblies, bobbins, winding interfaces, and fixtures. You will help drive motor development from initial design through repeatable low-volume builds while establishing the tooling, work instructions, and validation practices needed for future robotic platforms. This role is based in San Francisco, CA, and requires in-person presence 4 days a week. In this role, you will Own stator and rotor mechanical designs and released CAD and drawings, including geometry, interfaces, fits, tolerances, retention, assembly access, and mechanical validation. Develop laminations, stack assembly methods, bobbins, and winding interfaces that control alignment, insulation, conductor placement, and end-turn packaging. Design, fabricate, and debug fixtures for winding, stacking, assembly, alignment, and inspection. Build and troubleshoot prototypes. Use measurements, defects, rework, and assembly effort to improve designs and processes. Establish low-volume prototype production with equipment, build sequences, work instructions, revision control, traceability

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Employee Technology & Experience (ETX) team is responsible for delivering a world-class internal technology experience that enables employees to do their best work. We support and operate the employee-facing systems that keep the company moving quickly and efficiently. ETX spans support, logistics, AV, identity, endpoints, SaaS administration, automation, enterprise tooling, and internal infrastructure operations. We partner closely with Security, Engineering, Workplace, Finance, People, and other teams to keep OpenAI’s internal technology reliable, scalable, and moving at the pace of the company. About the Role We are hiring a Program Manager to help scale how IT operates across OpenAI. This role will lead complex cross-functional programs that improve operational maturity, streamline how teams work together, and turn high-impact initiatives into durable operational capabilities. You will work across IT, Security, Engineering, Workplace, and other functions to drive alignment, remove friction, and help build the operational foundation needed to support OpenAI’s rapid growth. You’ll be responsible for: Lead cross-functional operational programs that improve scalability, consistency, and operational maturity. Drive operational excellence initiatives across IT Support, employee lifecycle operations, meeting room and calendaring services, onsite support, vending, and research support environments. Build operating models, readiness plans, escalation paths, governance cadences, and success metrics for complex operational programs. Partner with technical teams to ensure new deployments, infrastructure investments, and internal platforms are operationally ready and sustainably supported at scale. Drive high-priority operational programs supporting company growth, including infrastructure expansion, operational integrations, and other emerging initiatives. Improve operational visibility, stakeholder alignment, and coordination across long-running cros

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We are looking for a systems-minded engineer to help advance our kernel development, performance engineering, and hardware-software co-design capabilities, with a particular focus on AI-assisted workflows and tooling. This person will work at the intersection of kernel optimization, developer tooling, observability, and research infrastructure, helping us improve both how production kernels are built and optimized, and how future hardware-software systems are designed and evaluated. The role is ideal for someone who is excited by low-level performance work, but also sees AI and automation as powerful tools for accelerating engineering velocity. You will help define the future of kernel engineering in the era of AI-assisted development. In this role, you may: Build developer tooling and workflows that make kernel development and performance optimization faster, more scalable, and easier to debug, integrate, and deploy. Develop observability, diagnostics, and validation infrastructure that makes AI-assisted optimization systems more interpretable, reliable, and effective. Optimize production kernels end to end by formulating optimization problems, running search loops, analyzing bottlenecks, debugging generated implementations, and landing improvements into production. Design abstractions, interfaces, and automation systems that accelerate kernel optimization, correctness validation, and hardware-software co-design. Improve AI-assisted optimization systems for sp

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a software engineer to help build the design methodology, software abstractions, and infrastructure that enable a small silicon team to develop complex chips rapidly and with high confidence. You will turn evolving architecture and design needs into reusable tools and workflows that improve iteration speed, quality, and then apply those tools to help construct world-class silicon. You’ll work closely across architecture, design, verification, performance modeling, and systems software. This role is well suited for an engineer who enjoys building high quality software and is motivated by the challenge of improving velocity and quality of the silicon development process. In this role, you will: Develop and scale design methodologies for rapid first-party chip development and apply them to construct complex custom chips Create abstractions that allow hardware structures, configurations, experiments, and results to be represented consistently across tools. Automate high-value engineering workflows and improve their reproducibility, observability, testability, and ease of use. Partner with architects, RTL designers, verification engineers, compiler engineers, and systems software engineers to gather requirements and then implement solutions. Use methodology and tooling to identify design risks early, accelerate iteration, and improve confidence in performance and implementation tradeoffs. Contribute across multiple aspects of software and hardware

PythonAWSGitRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling. We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI’s production and research workloads. About the Role On the Hardware Health and Observability team, you’ll build critical infrastructure that keeps OpenAI’s largest compute clusters healthy and operational at scale. Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams. Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally. In this role, you will: Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure. Build and evolve health checks that detect, remediate, and verify failures at scale. Ensure critical health checks execute with minimal latency to maximize workload uptime. Investigate hardware failures and system-level issues across large-scale compute environments. Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes. Build automation and tooling that enables global cluster management with minimal manual intervention. Partner with workload, reliability, and provider teams to integrate health signals into training and inference system

PythonSQLAWSLinux
P
📍 San Francisco, CA, United States· Full-time· Remote
✓ High-confidence listingCompany trend -85.6%

From $1.5M/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . Job Title: Software Engineer II, Data Analytics and Engineering Intro: We’re looking for a Software Engineer II, Data Analytics and Engineering to improve the quality, reliability and velocity of data science and product development at Pinterest. You’ll build scalable data foundations, analytics tooling and analysis pipelines that enable trusted, self-service access to datasets, insights and metric investigations across cross-functional teams. What you’ll do: Develop and document practical instrumentation and experimentation standards, then partner with product engineering teams to apply them to priority product development work. Build and improve scalable analysis pipelines and tooling that produce reliable insights at scale and strengthen understanding of key data structures and metrics. Create tools and processes that enable Data Scientists a

PythonSQLAWSRest
🔔

Get new aws and tooling platform lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime