We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. The Compensation team designs, manages, and evolves Plaid’s compensation strategy to attract, retain, and motivate world-class talent. We make sure compensation is fair inside the company and competitive in the market. We work on programs that connect directly to Plaid’s business goals and employee experience — using data and insight to make equitable, well-informed decisions. Our work spans cash and equity program design, annual and mid-year review cycles, benchmarking, and exec comp support. In this role, you will own compensation support for specific client groups and take the lead on cycle operations and comp tooling across the team. You’ll work alongside a sharp team and closely with Plaid’s Head of Total Rewards. You’ll have clear ownership of your domain from day one, with real exposure to equity program mechanics, exec comp, and board-level deliverables. Responsibilities: Own comp support for specific client groups — offers, promos, manager/HRBP questions — with a high bar for speed and accuracy Lead configuration and rollout of annual and mid-year compensation cycles in Workday and comp tooling Manage and maintain Plaid’s comp systems: Compa, Pave, and external survey tools (Compensia, Radf
Jobs in United States
Aws And Tooling Platform Lead in San Francisco
866 active opportunities · Updated October 2026
Showing
15 jobs
Explore current aws and tooling platform lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Privacy Engineering team builds secure, reliable systems that help OpenAI meet its legal obligations while protecting user data. We partner closely with Legal and Engineering teams across OpenAI to support lawful data access requests and other critical legal workflows. Our work turns complex, high-stakes processes into auditable and dependable technical systems with clear human oversight and strong privacy and security controls. About the Role We’re looking for a full-stack Software Engineer to build the internal tools and data pipelines that power lawful data access request workflows and Legal Operations. You will work across product and data systems to make authorized retrieval and case handling accurate, efficient, and auditable. This role is well suited to someone who enjoys translating ambiguous operational requirements into durable systems, cares deeply about sensitive-data handling, and wants to improve both technical reliability and the day-to-day experience of the people operating these workflows. This role is based in San Francisco, CA, with two additional locations under consideration: London, UK, and Dublin, Ireland. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and operate backend systems and workflow tooling for the full lifecycle of lawful data access requests, from intake and scoping through authorized retrieval, review, preparation, and audit. Build reliable data pipelines and interfaces across products and data stores so authorized teams can locate and handle the right records accurately and reproducibly. Implement least-privilege access, approval gates, provenance, audit trails, data minimization, and safe failure modes for sensitive workflows. Partner with Legal and Legal Operations to translate legal and operational requirements into clear technical designs and intuitive operator experiences. Identify responsible automation o
About the Team The Early Access Program (EAP) team leads high-impact alpha programs at the intersection of customers, Product, Research, Engineering, GTM, Security, Legal, and launch teams. We partner with customers to test emerging capabilities with real-world use cases, surface actionable insights, and support launch decisions. Our team is made up of builders who learn quickly, collaborate deeply, create clarity in ambiguity, communicate openly, and iterate constantly. About the Role We’re hiring an Early Access Deployment Engineer to lead technical engagements with customers who are leveraging our frontier capabilities to solve real-world use cases. You will work at the earliest—and often messiest—stage of development, when capabilities are still unclear, tooling and processes are evolving, and the path from promising technology to a valuable real-world application has yet to be defined. You will be a hands-on builder, problem solver, and technical partner to customers and our research/product team. You’ll push beyond initial assumptions, identify high-value use cases, prototype solutions, design useful evaluations, and troubleshoot what is and is not working. Managing multiple customer engagements at once, you will help customers navigate ambiguity and difficult technical decisions while translating their experience into clear, actionable feedback for Research and Product. You will also own the end-to-end execution of early access programs—from onboarding customers, supporting live experimentation, synthesizing findings, and informing launch decisions. You’ll collaborate deeply with Research, Product, Engineering, Applied Evals, GTM, Legal, Security, Marketing, and other launch partners to create clarity, manage risk, and keep programs moving through changing conditions. Success in this role means turning frontier capabilities into real-world customer value and high-quality research signals. You will develop reusable technical approaches from early deployments,
About the Team The Intelligence and Investigations team seeks to rapidly identify and mitigate abuse and strategic risks to ensure a safe online ecosystem. We are dedicated to identifying emerging abuse trends, analyzing risks, and working with our internal and external partners to implement effective mitigation strategies to protect against misuse. Our efforts contribute to OpenAI's overarching goal of developing AI that benefits humanity. The Strategic Intelligence & Analysis (SIA) team provides safety intelligence for OpenAI’s products by monitoring, analyzing, and forecasting real-world abuse, geopolitical risks, and strategic threats. Our work informs safety mitigations, product decisions, and partnerships, ensuring OpenAI’s tools are deployed securely and responsibly across critical sectors. About the Role As a Quantitative Intelligence Analyst , you will focus on discovering novel and emerging risks in complex human–AI systems before they are well-defined, measurable, or widely understood. You will use deep subject matter expertise and quantitative tooling to surface weak, early, and unconventional risk signals. You will build analytic models that explain how harms could emerge and translate ambiguous patterns into structured, data-driven insight. Your work will help identify potential gaps in policy or coverage and operationalize previously unmeasured problems into signals that can support detection, mitigation, and planning downstream. You will develop analytical frameworks that map how new risks form, evolve, and propagate as products change, policies shift, and external events unfold. Your analyses will directly inform strategic risk prioritization and planning across the company, with regular visibility through strategic risk products. This role is based in office (hybrid, 3 days/week). Relocation support is available In this role, you will: Discover and define new quantitative risk signals where no established metrics exist, using subject matter exp
About the Team The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software is reliable, testable, and ready to ship. We design and maintain automated test frameworks, hardware-in-the-loop labs, and release pipelines that keep quality signals trustworthy and enable rapid, safe product launches. Our work spans developer tools, automation, systems integration, and cross-team collaboration to ensure every release meets the highest standards. About the Role As a Software Engineer, Quality and Developer Tools , you will build and own the systems that validate our device software—from test frameworks and regression infrastructure to hardware-in-the-loop labs and release gates. You’ll design the tooling and automation that keep quality signals trustworthy, integrate them into CI/CD, and make it easy for engineers and QA vendor technicians to execute reliable, repeatable workflows. We’re looking for engineers with deep experience in software quality, automation, developer tooling, and hardware-software integration who thrive on building scalable, reliable systems for validation and release readiness. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Test infrastructure & frameworks: Design, implement, and maintain a unified test framework for device software across unit, integration, system, and end-to-end testing, with reproducible runs and integrations with GitHub, Linear, and Slack. CI/CD integration & releases: Integrate test suites with Buildkite, enforce promotion criteria for staging and production, auto-file regressions, and publish traceable artifacts and release notes. Hardware-in-the-loop lab design & orchestration: Plan and bring up racks, power and networking systems, and orchestration for device testing; support automated flashing, provisioning
About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re hiring a Developer Productivity Engineer to help scale the engineering systems, safeguards, and developer workflows that enable our teams to move quickly without compromising reliability or performance. This role sits at the intersection of developer experience, CI/CD infrastructure, release engineering, production readiness, and inference systems reliability. You’ll work on the tooling and operational foundations that support model launches, inference optimizations, cloud provider integrations, and large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous, high-ownership engineer who cares deeply about making other engineers faster, safer, and more confident. A major focus of this role will be improving the tooling and infrastructure around deploy gates for inference engine images. These systems help ensure that every image released to production and research is correct, numerically sound, free of regressions, and performant across key metrics like time-to-first-token (TTFT) and time-between-tokens (TBT). You’ll help harden the systems that catch issues before they reach production, reduce noise from flaky or infrastructure-related test failures, and improve automation around triage, ownership, debugging, and escalation when failures occur. You’ll also work on improving observability, rollout safety, release automation, and developer self-service tooling across a rapidly evolving inference stack. This is not generic internal tools work. The systems you build directly impact OpenAI’s ability to support new model launches, safely ship inference optimizations to the world, onboard new infrastructure providers, and operate one of the largest and most p
About the Team The Core Network Engineering team owns the end-to-end networking stack that connects OpenAI’s compute infrastructure — spanning global WAN/edge connectivity, data-center networking, and high-performance host/xPU networking used for large-scale training and inference workloads. This team is responsible for ensuring networking is never the bottleneck to model training efficiency, cluster reliability, or fleet expansion. They design and operate the systems that provide predictable, high-throughput, low-latency connectivity across some of the world’s most advanced AI infrastructure. About the Role We’re looking for engineers to help build and operate the networking foundation behind OpenAI’s frontier AI systems. Depending on your background and area of focus, you may work across host networking, datacenter fabrics, or global WAN infrastructure. The problems span low-level systems software, distributed infrastructure, protocol readiness, observability, performance engineering, automation, and large-scale network operations. You’ll work on systems where microseconds of latency, tail performance, and network reliability directly impact model training efficiency and production serving performance. This role is ideal for engineers who enjoy operating close to the hardware/software boundary and solving performance-critical infrastructure problems at massive scale. In this role, you will: Design, build, and operate networking systems that support large-scale AI training and inference infrastructure Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-perf
Employee Applicant Privacy Notice Who we are: Shape a brighter financial future with us. Together with our members, we’re changing the way people think about and interact with personal finance. We’re a next-generation financial services company and national bank using innovative, mobile-first technology to help our millions of members reach their goals. The industry is going through an unprecedented transformation, and we’re at the forefront. We’re proud to come to work every day knowing that what we do has a direct impact on people’s lives, with our core values guiding us every step of the way. Join us to invest in yourself, your career, and the financial world. About The role As a Vulnerability Management Engineer, you will support the identification, assessment, prioritization, and remediation of vulnerabilities across applications and infrastructure. Working under the guidance of senior team members, you will assist in understanding how vulnerable dependencies enter an application, identifying remediation options, and engaging with engineering teams to track fixes. You will contribute to the maintenance of internal vulnerability-management tools, such as scripts, documentation, and reporting. The ideal candidate will have a desire to grow their AppSec expertise, will be eager to learn about modern security tooling and automation, and will be comfortable using AI tools like Claude to assist with documentation, investigation, and scripting tasks while following company security and data-handling requirements. What you’ll do Perform regular vulnerability assessments using different tools. Regularly drive remediation and reporting of cataloged vulnerabilities. Assess discovered vulnerabilities and properly prioritize their scope, impact and necessary response actions. Conduct security reviews of our products and production infrastructure. Contribute to vulnerability management, application security and/or offensive/red-team operations. Engage in security audit
$295K – $380K/yr
About the Team The OpenAI Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings. Working across the entire model stack, we integrate cutting-edge hardware and software to explore a broad range of robotic form factors. We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives. About the Role As a Senior Software Engineer, ML Systems & Training Infrastructure, you will be a deeply hands-on engineering force multiplier for the robotics team. You will help keep the training framework and surrounding infrastructure healthy, review and improve code quickly, debug failures across ML systems and infrastructure, and unblock researchers and engineers when the path from idea to working training job gets rough. We’re looking for people who love writing, reading, reviewing, and fixing code; who can get productive quickly in unfamiliar systems; and who bring strong practical judgment without a lot of ego or process overhead. This role will be based in San Francisco, CA and be expected in office 5 days per week and offer relocation assistance to new employees. In this role, you will: Review, improve, and clean up code across training frameworks and adjacent infrastructure. Identify risky or low-quality changes before they land, and raise the code quality bar without slowing the team down. Debug issues across ML training systems, GPUs, clusters, networking, and related infrastructure. Help researchers and engineers unblock broken training jobs, flaky workflows, and brittle internal tooling. Improve the reliability, maintainability, and usability of the robotics team’s training framework. Move quickly on practical engineering problems that directly affect team velocity. You might thrive in this role if you: Have strong software engineering fundamentals and excellent code review judgment. Have experience with ML systems, training fr
About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Threat Intelligence team protects OpenAI’s technology, people, research, and infrastructure by proactively identifying and disrupting adversaries who seek to compromise our systems or misuse our models. We investigate sophisticated threats, build tooling to scale and augment analysis, and deliver intelligence that shapes security strategy and equips leadership with timely, risk-aware insights. We combine technical depth, investigative rigor, and strong cross-functional partnerships to uncover threats and drive impact across OpenAI’s security and research organizations. About the Role As a Technical Threat Investigator at OpenAI, you will help protect the company from sophisticated adversaries targeting OpenAI and the broader ecosystem, as well as those attempting to misuse our models in support of cyber operations. This is a deeply investigative role. You will independently conduct complex, end-to-end investigations into capable threat actors to understand their behavior, infrastructure, emerging techniques, and how AI is integrated into their workflows. You’ll use these insights to proactively identify malicious activity and drive detection, disruption, enforcement, and safety improvements across the company. You’ll translate your investigative findings into durable solutions that scale impact. You’ll build and own lightweight tooling, automate where it matters, and create AI-assisted workflows to make investigations faster, more repeatable, and more effective over time. In this role, you will: Conduct deep, end-to-end investigations into sophisticated threat actors interacting with OpenAI’s models, products, and broader ecosystem. Think like an adversary — model attacker behavior, anticipate misuse patterns, and proactively hunt for, identify, and disrupt malicious activity. Leverage internal telemetry, OSINT, vendor data, a
About the Team We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability Team Focus Areas Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely. Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads. Databases: Building high performance, distributed database systems that power all of OpenAI's product stack. In this
About the Role As a member of the Data team within the Go-to-Market organization, you will help build a data-driven culture, improve decision-making, and advance strategic initiatives through analytics. This is a full-stack data role spanning data modeling, metric definition, visualization, analysis, and self-service tooling. You will build trusted, scalable data sources and products that give the business reliable, actionable insights. The work calls for judgment: you will choose the tool, approach, and level of investment that best fit each problem, from a focused analysis to a durable production data product. As a core partner to the GTM organization, you will address both foundational and ad hoc analytics needs. You will turn complex data into clear narratives that help technical and non-technical audiences understand what is happening, why it matters, and what they should do next. In This Role, You Will Partner closely with GTM teams to proactively identify high-impact questions and translate business needs into data models, metrics, analyses, and scalable technical solutions. Define, source, validate, and operationalize the metrics that guide the business, helping teams incorporate them into planning and day-to-day decisions. Lead cross-functional data projects across established and emerging business areas, including setting the data strategy for greenfield domains. Build scalable data models and pipelines that integrate and transform data from multiple sources into trusted, accessible datasets. Create dashboards, reports, analytical tools, and other data products that enable stakeholders to answer questions independently. Own the lifecycle of metrics, analytical models, and data products from initial exploration and prototyping through production and ongoing maintenance. Choose the most effective approach for each problem—whether an analysis, metric, data model, visualization, or self-service product—based on the audience, urgency, complexity, and expected v
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. We are hiring for a leader for the newly forming Detection and Response team at Plaid. Our mission is to protect Plaid's financial infrastructure by detecting and responding to suspicious activity across the company. We are responsible for the entire lifecycle of detection and response, including detection infrastructure, AI triage and response, investigation tooling, Red Teaming, and Fraud Operations. Security is foundational to the trust thousands of businesses and millions of consumers place in Plaid, and we work directly to reduce risks and enable our business to move faster and safer. As the Head of Detection and Response, you will be the founding leader responsible for standing up and evaluating Plaid's detection and response team. You will lead a specialized technical team of analysts and engineers, gain deep experience partnering with the CISO and cross-functional engineering leaders, and build critical security infrastructure at scale. This is a unique leadership opportunity to manage a team that encompasses traditional security operations alongside Red Teaming and Fraud Operations, directly impacting Plaid's security posture and long-term stability. Responsibilities: Form and set up Plaid’
About the Team OpenAI's Industrial Compute organization is building and operating the infrastructure foundation for the next generation of AI. Infrastructure Operations works across facilities, hardware, network operations, incident management, data center engineering, delivery teams, and external partners to bring capacity online safely, understand its operational state, and improve it over time. As OpenAI's data center portfolio grows across first-party and partner-delivered capacity, the organization needs clear goals, trusted data, repeatable processes, and systems that make ownership, risk, readiness, and performance visible. This role will help build the operating mechanisms that allow Infrastructure Operations to scale with rigor. About the Role We are seeking a Technical Program Manager to own the systems, data, reporting, governance, and program-management backbone for Infrastructure Operations. Reporting to the Delivery & Operations Lead, you will translate strategy into executable goals and operating cadences, turn operational needs into software and data solutions, and create the mechanisms that keep a rapidly evolving organization aligned and accountable. This role will also own the current 1P+3P delivery-tracking layer within Operations: milestones, delivery timelines, quantity forecasts, risks, decisions, and executive reporting. You will partner closely with 1P Delivery Program Management, Compute TPMs, Data Center Engineering, construction, commissioning, and operations leaders to ensure that delivery information becomes complete, usable input for readiness, handover, and ongoing operations. You will own program health and the operating system around it: the goals, data definitions, workflows, reporting, decision paths, and follow-through that help functional DRIs execute. The ideal candidate is comfortable in ambiguity, technically fluent enough to implement real systems, and relentless about converting scattered information into durable mechan
Coder is looking for an experienced and detail-oriented IT Generalist to join our growing team. This is a full-time role ideal for someone who thrives in a fast-paced, hands-on environment. You’ll be the go-to person for user support, SaaS administration, and IT infrastructure, playing a key role in ensuring our team stays productive, secure, and well-equipped. You'll work closely with the larger IT team to keep things running smoothly - from onboarding new hires to managing devices and licenses to supporting SOC 2 compliance initiatives. If you're a strong communicator with a passion for IT operations and solving real problems for real people, this role is for you. This position follows a hybrid work model. Candidates should be local and able to work from our local office on a regular weekly basis, with scheduling flexibility. What you’ll do here Act as the first line of IT support for employees via Slack and our internal ticketing system Manage user onboarding/offboarding, including account provisioning through Okta and license management via our SaaS management systems Administer and maintain core tools like Google Workspace, Slack, Jamf, 1Password, and other business-critical SaaS apps Set up and manage macOS hardware inventory, including procurement, configuration, and asset tracking Support SOC 2 and GDPR compliance by following processes for access controls, audits, and data security Help scale IT operations as we grow - identify gaps, recommend tooling, and streamline support workflows Document processes and contribute to internal knowledge bases to improve self-service and transparency What we’re looking for Have 2-4 years of experience in IT support or systems administration Are proficient with Okta, Google Workspace, Slack, Zoom, and Apple device management (Jamf preferred) Are comfortable working independently in a remote environment and prioritizing across a variety of IT tasks Have a security-first mindset and understand the importance of access contro
Other cities to consider
More places hiring for this role
Get new aws and tooling platform lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime