Jobs in United States

Aws And Tooling Platform Lead in San Francisco

866 active opportunities · Updated October 2026

Explore current aws and tooling platform lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

P
📍 San Francisco, CA, United States· Full-time
✓ High-confidence listingCompany trend -85.6%

From $268.1K/yr

Quick readStrong listing-quality and freshness signals

About Pinterest: Millions of people around the world come to our platform to find creative ideas, dream about new possibilities and plan for memories that will last a lifetime. At Pinterest, we’re on a mission to bring everyone the inspiration to create a life they love, and that starts with the people behind the product. Discover a career where you ignite innovation for millions, transform passion into growth opportunities, celebrate each other’s unique experiences and embrace the flexibility to do your best work. Creating a career you love? It’s Possible. At Pinterest, AI isn't just a feature, it's a powerful partner that augments our creativity and amplifies our impact, and we’re looking for candidates who are excited to be a part of that. To get a complete picture of your experience and abilities, we’ll explore your foundational skills and how you collaborate with AI. Through our interview process, what matters most is that you can always explain your approach, showing us not just what you know, but how you think. You can read more about our AI interview philosophy and how we use AI in our recruiting process here . We are looking for a Sr. Staff Machine Learning Engineer to be the Technical Lead for the Content Quality who will build the overall technical strategy, unified technical architecture and define a roadmap for industry leading methodology. We are seeking strong hands on machine learning background including content modeling, signal lifecycle, and platforms used to enforce signal use with downstream use cases. You’ll be working with other leads to set and execute a long-term strategy for the team, aligning the strategy with other clients where it makes sense and communicating to leadership our current status and path to having world-class capabilities. You'll also foster a healthy community where all Content Quality engineers can learn best practices, collaborate effectively and understand our technical direction. What you’ll do: Arch

SQLAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Strategic Finance team partners across OpenAI to turn company strategy into financial decisions, helping allocate resources and steer the business toward its highest-impact long-term outcomes. Within Strategic Finance, the B2B Product Finance team owns the financial perspective on growth across OpenAI’s B2B portfolio. We connect monetization and customer economics with compute demand, capacity, margin, and resource planning, partnering closely with Product, GTM, Strategic Deal Desk, Compute Finance, Capacity Planning, Data, Accounting, and other Finance teams. About the Role We are hiring a Strategic Finance leader for B2B Product to build the 0→1 financial foundations and decision-making frameworks that will help OpenAI scale its B2B business sustainably. You will own high-impact work across B2B monetization, compute demand modeling, enterprise deal economics, contribution margin, and more. This is a portfolio-wide individual contributor role spanning our API Platform and other B2B products. You will connect customer demand and commercial terms to revenue, compute consumption, and margin outcomes; influence some of our largest enterprise deals; and help leadership make sound growth investment, compute capacity, and resource allocation decisions. This role is based in San Francisco, CA. In this role, you will: Own an integrated view of B2B monetization and product economics across the API Platform and other B2B products – including pricing, packaging, channel and usage mix, discounts, credits, commitments, revenue, and customer profitability Build, maintain, and improve complex driver-based financial models that connect customer usage, product and model mix, pricing, discounting, credits, commitments, and contribution margin Build financial views for B2B compute demand and margin planning; identify risks and opportunities, explain key drivers, and recommend capacity and resourcing decisions that improve portfolio economics Evaluate some of OpenAI’

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. Mitigating the frontier risks resulting from these capabilities is paramount to OpenAI’s ability to continue deploying models safely. The Preparedness team is dedicated to addressing these critical risks. Our work includes: Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. Mitigation. Keeping misuse and misalignment safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework , and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. About the Role We are seeking exceptional researchers who can push the frontier of safety mitigations. You will help derisk frontier models by developing novel safety mitigations, developing and applying new techniques from domains like interpretability, control, and alignment to ensure the safety of OpenAI’s deployed models. You will play a critical role in defining how a safe AI system should look in the future at OpenAI, making a significant impact on our mission to build and deploy safe AGI. This role requires strong technical depth and close cross-functional collaboration to ensure our safety mitigations are enforceable, scalable, and effective. We seek researchers who can partner with experts across domains such as misalignment, cybersecurity, and biology in order to develop the best possible end-to-end safety stack. In this role, you will: Work on identifying emerging AI safety risks and new methodologies for exploring and mitigating the impact of such risks Build (and then continuously refine) the evaluations that enable us to assess the extent of these risks; this might include worki

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team API Enterprise Controls is part of the API Infrastructure organization and owns the platform capabilities that help developers, startups, and enterprises adopt the OpenAI API securely and confidently. We build the systems underneath our APIs and developer platform across authentication and identity, service accounts and key management, secure networking, compliance, auditability, observability, and operational controls. Our users are developers and teams running critical applications on OpenAI, and we partner closely with Product, go-to-market, security, and infrastructure teams to turn their most important needs into reliable, intuitive platform capabilities. About the Role We are looking for an exceptional backend software engineer to help define and ship the enterprise capabilities our API Platform needs to scale.; this is a product-engineering role grounded in deep backend systems. You will work across databases, streaming systems, request routing, authentication, and developer-facing APIs while bringing strong product judgment, developer empathy, and attention to the small details that make a platform easier to understand, trust, and operate. You will lead large cross-functional initiatives, work closely with Product and go-to-market teams, engage directly with sophisticated users, and carry ambiguous needs from discovery through design, launch, and iteration. In this role, you will: Own backend product capabilities end to end across authentication and identity, service accounts and key controls, secure networking, compliance, observability, and operational workflows. Partner with Product, go-to-market, security, infrastructure teams, and sophisticated customers to identify needs, shape the roadmap, and lead large cross-functional projects from design through launch. Design developer-facing APIs, system behavior, configuration, error handling, safe defaults, auditing, and notifications with exceptional care for the details that define a great dev

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Professional Services team helps organizations move from AI ambition to durable production outcomes. We partner with customers on complex deployments and build the strategy, commercial models, operating mechanisms, and delivery capacity needed to realize value from OpenAI's models and products. The team works across Go-to-Market, Forward Deployed Engineering, Technical Success, Finance, Product, Legal, Data, Systems, and delivery partners. We are building a services motion that is customer-centered, commercially rigorous, operationally scalable, and deliberately connected to product adoption. About the Role We are seeking a senior GTM Strategy & Operations professional to build and scale the operating system for OpenAI's Professional Services business. This is a foundational, hands-on individual contributor role at the intersection of business strategy, finance, go-to-market, and delivery. You will turn ambiguous questions—what we offer, how we price and package it, how we plan capacity, and how we measure performance—into clear decisions and repeatable mechanisms. You will own business planning, pricing and packaging, forecasting and modeling, management reporting, and cross-functional strategic initiatives. You will build integrated views of demand, staffing, revenue, margin, and delivery performance; replace one-off analyses with durable processes; and create operating cadences that help leaders act early. You will be a trusted partner to Professional Services, GTM, FDE, Technical Success, Finance, Product, Legal, Revenue Operations, and Data leaders. The right person combines direct Professional Services judgment with rigorous analytics, executive communication, and the willingness to build the model, process, or dashboard themselves. You’ll be responsible for: Define and drive the business and GTM strategy for Professional Services, including target customer needs, offer portfolio, positioning, pricing and packaging, partner motions,

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Our Cyber team builds AI systems and products that help trusted defenders understand and respond to cyber threats while improving the safety and reliability of frontier models in security-sensitive settings. The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced cyber capabilities useful to defenders and responsibly managed. We collaborate closely with Safety/Preparedness, Research, Security, Legal, Communications, GTM, and external partners across OpenAI’s broader cyber work. About the Role We’re looking for research and software engineers to join Codex Cyber. You’ll help define and ship security products, work with trusted defenders and customers, shape model training and access patterns, and build research and evaluation systems for assessing cyber capabilities, validating safeguards, and improving training data. This role is hands-on and cross-functional, connecting product launches, model development, safety work, and real-world security use cases. In this role, you will: Help define and execute the technical roadmap for Codex Cyber’s security products, including evaluations, safeguards, trusted-defender workflows, and deployment decisions. Work with trusted defenders, customers, and partner teams to understand cyber use cases, evaluate risk, and turn feedback into product and research priorities. Shape cyber-specific model training and access patterns, including data, evaluations, validation, and deployment criteria. Build and validate systems for measuring cyber capabilities, monitoring misuse risk, and proving safeguards work in practice. Collaborate with Safety/Preparedness, Research, Security, Legal, Communications, Go-to-Market, and external partners on company-wide cyber priorities. Translate frontier cyber research into launch-ready tools, operational playbooks, and durable infrastructure for Codex and security products. You might thrive in this role if you: Enjoy 0 -> 1 envi

JavaScriptTypeScriptPythonJava
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads. About the Role We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network. The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis. You will partner closely with network architecture, systems engineering, GPU engineering, storage engineering, security, deployment, site operations, service providers, colocation partners, and hardware vendors to raise reliability and reduce operational toil. Key Responsibilities Own the operational health, availability, and reliability of production AI network infrastructure across Industrial Compute's data centers. Monitor, troubleshoot, and resolve network incidents while meeting service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR). Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks. Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact. Manage the hardware lifecycle, including switch and optics replacements, RMA coordination, software upgrades, and preventive maintenance. Support new A

PythonAWSAzureGit
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's Marketing team helps customers understand, adopt, and get value from OpenAI's products and platform. The SMB Ads Marketing team is building the growth engine for small and medium-sized business advertisers: helping them discover the platform, sign up, launch campaigns, understand performance, and grow spend over time. We take a data-driven approach to understanding customer needs, market dynamics, and funnel behavior. We partner closely with Product, Engineering, Data Science, Sales, Partnerships, Marketing Operations, Design, and Communications to create a cohesive end-to-end advertiser experience and help businesses unlock value from OpenAI's advertising platform. About the Role We are hiring a Growth Marketing Manager, SMB Ads to help build and execute high-velocity growth programs for small business advertisers. This person will work across acquisition, activation, lifecycle, and conversion, with an initial focus on helping advertisers move from interest and sign-up to first meaningful spend & retention over time This is a hands-on growth role for someone who loves turning ideas into shipped experiments. You will build campaigns, lifecycle journeys, landing page tests, messaging tests, audience experiments, AI-assisted workflows, reporting improvements, and practical conversion programs. As the SMB Ads business grows, this role may continue to support acquisition and activation and expand deeper into same-store growth, lifecycle, retention, and repeat spend. In This Role, You Will: Execute growth experiments across acquisition, activation, lifecycle, and early retention. Build email, in-product, landing page, audience, messaging, and offer tests that help SMB advertisers reach first spend and repeat value. Use AI tools to prototype marketing assets, automate workflows, analyze funnel data, generate creative variations, and build internal growth systems. Partner with Growth, Product, Engineering, Data Science, and Marketing Operations

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve. Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes. We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams. Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products). About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas. This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers. In all the projects this role pursues, the ultimate goal is to push the field forward. We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code. Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact. This role is based in San Francisco, CA. We use a

PythonAWSRestMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) which benefits all of humanity. This long-term undertaking brings the world’s best scientists, engineers, and business professionals into one lab together to accomplish this. In pursuit of this mission, our Go To Market (GTM) team is responsible for helping customers learn how to leverage and deploy our highly capable AI products across their business. The team is made of Sales, Solutions, Support, Marketing, and Partnership professionals that work together to create valuable solutions that will help bring AI to as many users as possible. About the Role Our GTM team is uniquely positioned to help customers realize the transformative potential of advanced AI models for their businesses and end users. As an individual contributor on the GTM Operations team, you’ll play a critical role in designing and scaling the operational systems that power our sales organization. This role will serve as a trusted partner to GTM leadership, building the end-to-end ops design for sales lifecycle from lead routing through territory design, opportunity management, deal execution, and delivery readiness. This role combines systems and process design with operational performance management, delivering insights and driving automation to improve field efficiency and velocity. You’ll collaborate cross-functionally with Marketing Ops, Enterprise Systems, Product, Delivery, Finance, Enablement, Legal, Deal Desk, and Security to develop scalable infrastructure, streamline workflows, and enable scalable growth across the business. In this role, you will: GTM Data,Governance & Routing: Create a reliable GTM data foundation that makes SFDC easier to use and ensures leads, accounts, and opportunities are accurately routed, defined, enriched, and actionable. Design and manage lead and campaign routing; define requirements and partner with systems and marketing ops on build. Implement alerting, monitoring, an

SQLAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the team The Applied team safely brings OpenAI's technology to the world. We released ChatGPT; Plugins; DALL·E; and the APIs for GPT-5, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon. Our customers build fast-growing businesses around our APIs, which power product features that were never before possible. ChatGPT is a prime example of what is currently possible. We simultaneously ensure that our powerful tools are used responsibly. Safe deployment is more important to us than unfettered growth. The Fraud Engineering team works within our Applied Engineering organization identifying and responding to fraudsters on our platform. We are looking for a software engineer with anti fraud & abuse experience to help architect and build our next-generation anti-fraud systems. About the role The Scaled Abuse team protects OpenAI’s products and customers by detecting, preventing, and responding to fraudulent and abusive behavior at scale. We build and operate the backend and data systems that power real-time detection, investigation workflows, and enforcement — balancing strong protections with a great user experience as the platform grows. Our work sits at the intersection of engineering and abuse expertise: we partner closely with Trust & Safety, Security, and Product to understand emerging attack patterns, translate messy signals into clear system behavior, and continuously harden our defenses. The problems are dynamic and ambiguous by default, so we value engineers who can quickly dive into an unfamiliar codebase, develop strong intuition about how it works end-to-end, and propose pragmatic improvements that make the entire stack more resilient. In this role, you will: Design and build systems for fraud detection and remediation while balancing fraud loss, cost of implementation, and customer experience Work closely with finance, security, product, research, and trust & safety ope

PythonAWSAzureKubernetes
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects, and spend less time on repetitive tasks. AI has already transformed how code is written, but software engineering extends far beyond coding. Our mission is to apply AI across the entire software development lifecycle (SDLC) — from design and implementation to code review, testing, debugging, issue remediation, maintenance, documentation, and user support. The team is also responsible for developer-facing Codex experiences including the Codex IDE Extension and the terminal interface, which are used daily by developers ranging from individual open-source contributors to some of the world’s largest engineering organizations. The team also works closely with the open-source software community, building tools that help maintainers and contributors manage increasingly complex projects. We believe AI can make open-source development more sustainable by reducing the operational burden of reviewing contributions, triaging issues, maintaining quality, and supporting growing communities. By building the future of software development, we're helping advance OpenAI's mission of ensuring that the benefits of AI reach people around the world. About the Role We’re hiring a Full Stack Software Engineer to help invent the next generation of AI-powered software development workflows. “Full stack” in this role means much more than traditional frontend and backend development. You'll own complete product experiences, spanning user interfaces, workflow orchestration, agent and prompt design, backend systems, and cloud infrastructure. This is a highly product-oriented role. You'll work directly on the workflows developers use every day, identifying bottlenecks and rethinking how software gets built in a world where AI agents are active participants in the development process. The features you ship will inf

TypeScriptAWSRestAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity. The Security team protects OpenAI’s technology, people, and products. We are technical in what we build but are operational in how we do our work, and are committed to supporting all products and research at OpenAI. Our Security team tenets include: prioritizing for impact, enabling researchers, preparing for future transformative technologies, and engaging a robust security culture. About the Role We are seeking a Software Engineer, Security Observability to join our Security team. In this role, you will be responsible for building secure, scalable systems that enhance our security observability infrastructure. Leveraging your strong engineering skills, you will collaborate with cross-functional teams to develop, deploy, and maintain robust software solutions that support our security and detection capabilities. This role is open to remote employees, or relocation assistance is available to one of our OpenAI offices in San Francisco, Seattle, or New York City. Due to requirements associated with work this role may support, applicants for this position must be U.S. citizens. In this role, you will: Design and develop scalable software systems that facilitate security observability across our infrastructure. Build and maintain data pipelines that centralize and store security-relevant data from diverse sources. Proactively improve the resilience and reliability of data systems to ensure high platform availability Collaborate closely with Detection & Response (D&R) and other security teams to reduce the company’s security risk. Contribute to data engineering in support of forensic investigations and compliance efforts. You might thrive in this role if you have: Strong software engineering experience, with proficiency in programming languages such as Python, Golang, or similar. A background in infrastructure as code, with exp

PythonAWSAzureRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Codex Web Layer team provides the web-based systems and user experiences for Codex across the entire stack, from the Electron-like application framework that powers the application, to the user-facing in-app browser. About the Role In this role, you will be responsible for designing and implementing infrastructure and features end-to-end for the Codex desktop client application. You will help define what it means to be a hybrid agentic/interactive web browser. The team embodies “full stack” development from the lowest-level OS integration to the highest-level interaction design. This role is based in San Francisco. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role you will: Partner closely with product and design to conceive, design, and build features for Codex web browsing features on macOS and Windows. This role will focus mostly on the backend C++ layer and Chromium, but many features cross the full stack including some TypeScript. Partner with the wider Codex team to deliver a high-performance, stable, and secure application platform for client development. This includes API design and implementation (mostly in C++) and the infrastructure that supports deploying it (in Python, TypeScript, and agentic skills). Work with a small, experienced team of engineers on this critical and rapidly growing product. You might thrive in this role if you: Have significant experience building technically complex features end-to-end. Are a strong C++ developer, especially with experience in browser environments like Chromium and Electron. Since this role is more backend focused, general knowledge of web development and TypeScript is helpful but not required. Thrive in a fast-paced, ambiguous environment. Communicate clearly and concisely across many different roles in the organization. Are self-directed, identifying important work and executing it end-to-end. About OpenAI OpenAI is an AI

TypeScriptPythonAWSRest
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team OpenAI's mission is to ensure that artificial general intelligence benefits all of humanity. The Consumer Devices team is building a new generation of AI-powered products that seamlessly integrate hardware and software to create intuitive, transformative experiences. We bring together experts across embedded systems, machine learning, hardware, design, and product engineering to develop products at the intersection of AI and consumer technology. About the Role OpenAI is seeking a System Power Engineer to characterize, measure, and optimize power consumption across our embedded hardware products. In this role, you will work closely with Electrical Engineering and system software teams to build power test automation, measure subsystem-level power usage, and drive improvements that directly impact battery life, thermal behavior, charging performance, and system reliability. You will help establish the methodologies and metrics used to understand and improve power efficiency across real-world product experiences, from controlled lab environments to representative day-in-the-life usage scenarios. This role requires hands-on experience with embedded hardware platforms, power instrumentation, and the analysis of power profiles and system behavior. This role is based in San Francisco, CA. We use a hybrid work model of four days per week in the office and one day working remotely. Relocation assistance is available for new hires. In this role, you will: Define and develop power testing automation to evaluate system behavior across a range of workloads and operating conditions. Measure subsystem-level power consumption using power breakout probes and other lab instrumentation. Develop and execute power characterization tests spanning basic workloads, complex mixed-use scenarios, and representative day-of-use experiences. Partner closely with Electrical Engineers to identify opportunities to improve system power efficiency. Collaborate with software engineering

PythonAWSRestMachine Learning
🔔

Get new aws and tooling platform lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime