About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters. Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads. You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems. This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet. This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office). Relocation assistance is available. In this role, you will: Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads. Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware. Engineer and optimize compute and data kernels, ensuring correctness, high performance, and portability across simulation and production environments. Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales. Develop simulation infrastructure to validate runtime b
Jobs in United States
Aws And Tooling Platform Lead in San Francisco
866 active opportunities · Updated October 2026
Showing
15 jobs
Explore current aws and tooling platform lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.
About the Team The Intelligence and Investigations team seeks to rapidly detect and disrupt abuse in AI technologies to ensure their safe use. We are dedicated to identifying emerging abuse trends, analyzing risks, and working with our internal partners to implement effective mitigation strategies to protect against misuse. Our efforts contribute to OpenAI's overarching goal of developing AI that benefits all of humanity. About the Role As an Intelligence Systems Engineer, you’ll be focused on advancing our Intelligence & Investigations efforts at OpenAI, ensuring the safe and responsible use of AI across our products and services. We are seeking a self-starter to prototype, develop, and maintain new tools and processes that integrate OpenAI’s models and infrastructure to enable internal teams to make sense of large, open-domain datasets, fight abuse, and inform high-stakes decisions. You will be a crucial technical bridge between our data scientists and subject matter experts and technical teams like Platform Integrity, Safety Systems, and Research by leading the development of innovative tools and processes that bolster goals in scaled collections, investigations, and analysis. The ideal candidate has strong analytical and data skills, with a background in both prototyping and building scalable systems that can swiftly detect emerging threats, process vast amounts of information, and deliver insights to stakeholders. We value professionals with outstanding communication skills, a commitment to continuous learning, and who are dedicated to promoting the responsible use of AI. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Prototype, build, and maintain at-scale intelligence systems that detect, triage, and monitor targeted signals from both open-source and internal data Analyze requirements and deliver end-to-end solutions that address
About the Team The Core Services team is responsible for building and managing foundational services. It acts as the bridge between core infrastructure (e.g. compute, storage, networking) and product engineering teams, and enables product teams to move fast, build reliably, and scale efficiently. About the Role As a software engineer in the core services team, you will design and operate critical backend platforms such as caching systems, workflow orchestration, metadata stores, and file services. You’ll focus on building highly reliable, scalable, and performant systems that serve as the backbone of our products. We’re looking for people who are passionate about building infrastructure that empowers product teams, love working on distributed systems challenges, and enjoy creating well-designed APIs and abstractions that accelerate development. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Design, build, and maintain shared infrastructure services such as caching layers, workflow orchestration (Temporal), metadata stores, and file storage services. Collaborate with product teams to provide scalable, reliable primitives that abstract the complexities of distributed systems. Improve performance, resilience, and scalability of core services that power customer-facing applications. You might thrive in this role if you: Have experience with distributed systems, caching infrastructure (e.g., Redis, Memcached), metadata storage (e.g., FoundationDB), or workflow orchestration (e.g., Temporal, Cadence). Have experience running containerized services in cloud environments and integrating them into automated build/test/release (CI/CD) workflows. Understand trade-offs in consistency models, replication strategies, and performance optimization in multi-region systems. Excel at communication and collaboration with cross-functional teams, and are obsesse
About the Team We are building general-purpose robotics. In the short term, we are focused on robots to support skilled workers to build our future infrastructure. In the long term, we imagine everyone having a personal robot doing anything they need. Progress is rapid, and based on a foundation of co-design between robotics hardware and ML research. About the Role As a Firmware Engineer, you will define and drive the architecture of embedded systems for next-generation hardware products. You will own foundational firmware decisions across real-time execution, device bring-up, hardware interfaces, fault handling, safety mechanisms, and production readiness. We’re looking for someone with deep experience building safety-critical or high-consequence systems, where failures can have meaningful consequences. You should be comfortable reasoning about risk, designing for diagnosability and graceful degradation, and creating engineering practices that raise the reliability bar for the entire team. You should also be unusually good at moving fast. Sometimes the right answer is a carefully reviewed architecture that will endure for years; sometimes it is getting a rough-but-useful prototype working by the end of the afternoon so the team can learn something concrete tomorrow. We value engineers who know the difference, make that call well, and can operate credibly in both modes. You will be both a technical leader and a hands-on builder: setting direction, reviewing critical designs, unblocking the hardest problems, and writing production firmware when it matters most. Our embedded stack uses a lot of Rust. Extensive experience in the language is a big help! This role is based in San Francisco, CA. This role will be expected to be in office 4 days per week and offer relocation assistance to new employees. In this role, you will: Rapidly bring up new hardware and set execution pace for the team. Lead firmware architecture for embedded systems spanning boot, RTOS/runtime behav
About the Team The Software Engineering team is responsible for designing and building the scalable, performant, and secure backend systems that power our products—from early prototypes to large-scale deployments. We collaborate closely with product, hardware, and full-stack teams to ensure our infrastructure enables fast iteration while setting a strong foundation for long-term growth. About the Role As a Backend Engineer , you will design and build services, APIs, and infrastructure that support evolving product needs. You’ll apply a deep understanding of backend systems and maintain enough end-to-end context—from hardware to cloud—to guide technical decisions that best serve the product and team. We’re looking for engineers who thrive in fast-paced, collaborative environments and care deeply about building robust systems that scale. This role is based in San Francisco, CA . We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Architect, build, and maintain high-performance, secure backend systems. Design APIs, data models, and infrastructure to support evolving product needs. Balance near-term development velocity with long-term maintainability and scalability. Collaborate with cross-functional teams to ensure cohesive, end-to-end solutions. You might thrive in this role if you: Have 7+ years of professional software engineering experience, with a focus on backend systems. Have a proven track record of building and scaling systems from early stage to large scale. Are proficient with Python and Go, and familiar with a range of server-side technologies. Have a strong grasp of system design, performance optimization, and security best practices. Can reason about full-stack tradeoffs from hardware through cloud infrastructure. (Nice to have) Have experience with distributed systems and cloud architectures. (Nice to have) Bring a background in instrumentation, analytics, and performanc
About the Team OpenAI’s Applications Engineering organization builds and operates the products that bring our cutting-edge research to millions of users and developers worldwide. The Applied Foundations team owns the core product and platform layers that make those experiences possible — from identity & access, to safety to payments & commerce across all of our apps. Our teams span product engineering, infrastructure, and safety, working together to deliver technology that is reliable, secure, and trusted at global scale. About the Role You will be a Senior Android engineer on OpenAI’s Applied Foundations team, building the core mobile experiences that power how users sign up, manage their account, family features, pay for services, stay safe, and interact with OpenAI’s products with confidence. This role is about creating high-quality products as well as reusable Android foundations that product teams across different OpenAI apps depend on to ship quickly while meeting the highest standards for security, reliability, and user trust. You’ll own complex client-side systems spanning UI, networking, local state, payment integrations and Apple platform integrations, and work closely with backend, product, and safety partners to shape the architecture that supports OpenAI’s mobile ecosystem at global scale. You might thrive in this role if you: Have 4+ years of professional software engineering experience. Have a proven track record of building high-quality Android applications in production. Are fluent in Kotlin (and/or Java) and familiar with Android development tools and architecture components. Prioritize performance, security, and user experience in mobile development. Enjoy working cross-functionally to bring ambitious product ideas to life. Care deeply about performance, security, and user experience. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities Co-design future hardware for programmability and performance with our hardware vendors Assist hardware vendors in developing optimal kernels and add support for it in our compiler Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory h
About the Team The Release Engineer team is responsible for building and maintaining the systems that power software delivery—from CI/CD pipelines and artifact management to release automation and fleet telemetry. We ensure software across bootloaders, firmware, operating systems, and cloud services is built reproducibly, validated rigorously, and released safely at scale. About the Role As a Release Engineer, you’ll design, build, and operate release infrastructure that enables reliable, secure, and traceable software delivery across complex multi-component systems. You’ll partner closely with embedded, cloud, and QA teams to ensure that every build—from development to OTA deployment—is fast, verifiable, and production-ready. We’re looking for engineers who take pride in automation, build reproducibility, and system reliability—and who enjoy building the connective tissue that allows hardware and software to ship together seamlessly. This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees. In this role, you will: Design and operate CI/CD pipelines for multi-component builds (bootloader, firmware, OS images, backend, companion apps) using hermetic toolchains. Define versioning and branching strategies; automate promotions, changelogs, and artifact retention. Integrate unit, integration, and hardware-in-the-loop (HIL) test results; quarantine flaky tests, auto-bisect failures, and block unsafe promotions. Build A/B OTA update flows with verity and health checks; run staged rollouts and canaries; implement safe rollback and roll-forward strategies. Implement code signing for binaries and firmware, generate SBOMs, run vulnerability scanning, and attach build attestations and provenance. Manage dashboards and alerts for build health, promotion latency, failure rates, and fleet update telemetry. You might thrive in this role if you: Have experience building and operating buil
About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing the world’s most exceptional talent under one roof to push the boundaries of what’s possible. Our Research Recruiting team plays a critical role in this effort. We are an embedded part of the research organization, working side by side with our research staff to deeply understand evolving priorities, build trust, and strategically shape the future of OpenAI’s talent. About the Role You will own and execute long-term talent strategies to identify, engage, and recruit many of the world’s leading and emerging AI researchers, research engineers, and technical scientists working at the frontier of machine learning. This is not a traditional execution-focused recruiting role. You will operate as a strategic partner to OpenAI’s research staff, helping define hiring priorities, shape search strategy, influence candidate evaluation, and guide hiring decisions that directly impact the direction and quality of our frontier-model research and fulfillment of our mission. In this role, you will: Partner directly with research and technical staff to define hiring priorities, shape search strategies, and anticipate future talent needs as technical roadmaps evolve. Proactively identify and cultivate exceptional AI/ML research talent across industry, academia, and emerging labs, often before formal hiring needs exist. Use market insights and candidate signals to influence hiring decisions, leveling, and compensation strategy for highly specialized research roles. Serve as a trusted advisor throughout candidate evaluation and closing — helping leaders calibrate for research excellence, long-term potential, and organizational fit. Collaborate closely with your sourcing partner to execute complex, high-impact searches in ambiguous or rapidly evolving technical domains. You might thrive in this role if you: Significant experience recruitin
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models. Our work spans system software, networking, platform architecture, fleet-level monitoring, and performance optimization. About the Role We’re hiring an SW Engineer to enable production workloads and end-to-end testing on new platforms. This role will include creating new test harnesses and platform stress benchmarks, porting existing inference and training workloads to new, sometimes early-access, systems/hardware, analyzing performance and bottlenecks, and characterizing the end-to-end behavior of new systems (compute, comms, storage, control plane, and failure modes). Key Responsibilities Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar. Build a suite of benchmarks and stress tests that capture real E2E behavior of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts. Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries) Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects Create repeatable test harnesses that run in CI / lab environments and produce actionable outputs (pass/fail, performance score, regression detection). Partner with systems + fleet bring-up engineers to ensure the platform is not only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops). Work cross-functionally with vendors and internal stakeholders by producing
About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical. The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly. About the Role We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity. In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity. This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions. This role is based in San Francisco, CA, with travel as needed. Key Responsibilities Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint Drive readiness to convert contracted compute capacity into schedulable production clusters Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives Build integrated schedules spanning procurement, logistics, installation, st
About the Team The Strategic Finance team provides financial insights and guidance to support OpenAI’s long-term goals and strategies. We partner across the business to allocate and deploy our resources for the highest-impact outcomes.  Within Strategic Finance, the B2B team focuses on the financial performance of our products and GTM functions, ensuring tight alignment between financial objectives and company strategy. We partner with leaders across Product, GTM, Research, Partnerships, and Operations to: Drive operational planning, financial forecasting, and performance management. Provide decision-quality insights on product and financial performance to inform strategic resource allocation. Build the “0→1” financial foundations required to scale and accelerate growth. About the Role We are hiring a senior leader in B2B Strategic Finance to build and scale a new pillar within our finance organization. This is a highly visible role that reports into the Head of B2B Strategic Finance and supports some of our most critical executive stakeholders, including our COO, CFO, and CRO, among others on the B2B Leadership Team. This role is ideally based in our San Francisco HQ, but we are open to NYC and Seattle. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Drive B2B finance scale and rigor Build and scale core financial infrastructure across the B2B business, including forecasting methodology, variance management, and performance narratives that drive accountability and decision-making velocity. Lead consolidated planning across revenue, gross margin (including compute), and opex for annual budget, forecasts, regular business reviews, and long-range planning. Establish durable management reporting: KPI definitions, dashboards, month-end/quarter-end deliverables, and exec-ready readouts. Partner with Corporate FP&A, Accounting, and Finance Systems/Data to evolve processes and contro
About the team OpenAI’s mission is to ensure the responsible and widespread adoption of artificial intelligence. In support of that mission, the Sales team partners closely with customers to deeply understand their businesses and needs, helping inform the development of products and solutions that drive meaningful revenue growth and long-term success on the platform—while maintaining strong standards for user trust and platform integrity. About the role We’re looking for an Account Manager to support advertisers in onboarding, launching, and growing on OpenAI’s advertising platform. You will manage a portfolio of partners and serve as a trusted advisor, helping them achieve their business objectives while ensuring a strong and responsible experience on our platform. This role sits at the intersection of relationship management, strategy, and execution. You’ll work closely with Sales to drive long-term partner growth and with internal teams to continuously improve the advertiser experience. In this role, you will: Own relationships with a portfolio of advertising partners, serving as their primary point of contact. Guide advertisers through onboarding and campaign launches to ensure strong adoption of OpenAI’s ads products. Analyze campaign performance and provide clear, actionable recommendations to improve outcomes. Identify opportunities to expand partnerships through increased product adoption and incremental investment. Partner cross-functionally with Sales, Product, Analytics, Policy, and Operations to resolve issues and enhance platform value. Surface advertiser feedback and insights to inform product improvements and long-term strategy. Help build scalable processes and best practices that support the growth of OpenAI’s advertising ecosystem. You might thrive in this role if you: Has 5-10+ years of experience in digital advertising, account management, customer success, consulting or a related client-facing role. Has experience managing advertiser or agency r
About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing the world’s most exceptional talent under one roof to push the boundaries of what’s possible. Our Research Recruiting team plays a critical role in this effort. We are an embedded part of the research organization—working side by side with our research staff to deeply understand evolving priorities, build trust, and strategically shape the future of OpenAI’s talent. About the Role You’ll focus on identifying and engaging world‑class AI researchers and scientists who are advancing the frontier of machine learning and artificial intelligence. In this role, you will drive top‑of‑funnel sourcing efforts, build thoughtful outreach strategies, and cultivate long‑term relationships with highly specialized talent across academia and industry to support our Research teams’ hiring ambitions. You’ll work across a range of deeply technical and research‑oriented roles, mapping emerging fields, identifying leading contributors, and proactively engaging experts in their respective domains. You’ll partner closely with recruiters, hiring managers, and research leaders to surface exceptional candidates, generate strong pipelines for niche searches, and continuously refine sourcing strategies to help our research teams grow and scale. Your Responsibilities: Build highly targeted searches that deliver candidate profiles with precision. Craft thoughtful approaches to candidate outreach that will engage even the most passive talent and cultivate relationships with high priority targets over the long-term. Develop and utilize creative strategies to discover non-traditional, research-adjacent, and/or up and coming talent. Partner closely with hiring managers across OpenAI to develop customized sourcing strategies and advise them on the market. Take a highly organized and data driven approach to candidate tracking and funnel metrics. Represen
About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. Make a Difference: Monitor and maintain deployed m
Other cities to consider
More places hiring for this role
Get new aws and tooling platform lead jobs in San Francisco, United States by email
Daily job updates · Unsubscribe anytime