Jobs in United States

Aws And Tooling Platform Lead in San Francisco

866 active opportunities · Updated October 2026

Explore current aws and tooling platform lead jobs in San Francisco. Filter by work mode, employment type, experience, department, date posted and distance.

C
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -79.2%

£225K – £325K/yr

Quick readStrong listing-quality and freshness signals

Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! As a Manager of Security Engineering, your key responsibilities include: Serve as trusted advisor to team’s leadership and partner teams by clearly articulating business risks associated with security issues Execute the long-term vision for the Security team in alignment with Cohere’s product and business goals. Collaborate closely with leadership to prioritize high-impact initiatives and strategic customer engagements. Vulnerability Management: Develop and implement enterprise-wide vulnerability management processes and tooling, including identification, prioritization, remediation tracking, and reporting, including customer artifacts Static Application Security Testing (SAST): Establish SAST programs, integrate tools into CI/CD pipelines, and analyze results to identify and remediate security flaws in source code Dynamic Application Security Testing (DAST): Implement DAST methodologies, configure scanning tools, and conduct regular assessments of running applications Penetration Testing: Lead and oversee internal and external penetration testing engagements, including web application, API, network and agentic AI platform inclu

PythonAWSAzureGCP
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Growth team drives user and revenue growth across ChatGPT’s consumer and business segments as well as other OpenAI products worldwide. We operate across the full funnel - from awareness and acquisition through activation, retention, and expansion - using a combination of global performance marketing, AI-powered workflows, in-product optimization, insights, experimentation, and creative ops engineering. About the Role We are hiring a Lifecycle Lead to build the company-wide owned-channel capability that helps teams reach users with relevant, timely, and trustworthy experiences. This senior, hands-on leader will set the lifecycle strategy, partner with Engineering to build the orchestration and deployment platform, and establish the operating model that allows teams across the company to launch and improve evergreen programs safely at scale. You will sit at the intersection of platform, product, and campaign strategy. You will partner with Engineering, Product, Data Science, and Analytics on the underlying systems, and with Product Marketing Managers and other client teams to design journeys that help new, active, and returning users reach value and build durable habits. In this role, you will: Partner with product to set the company-wide vision, roadmap, and operating model for lifecycle and owned-channel engagement. Partner with Engineering, Product, Data Science, and Analytics to shape the tooling and infrastructure for identity, audiences, eligibility, consent, triggers, orchestration, decisioning, frequency, experimentation, localization, quality assurance, and observability. Define scalable deployment workflows—including self-service and centrally supported paths, intake, templates, approvals, governance, service levels, and incident response—so teams across the company can launch safely and efficiently. Partner with Product Marketing Managers and other client teams to translate audience, product, and business goals into evergreen journey stra

AWSRestAIGo
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Team: Plaid is becoming an AI-first company, and Intelligent Tooling builds internal platforms and tools to lead the transformation. Our biggest opportunity isn't just better tools for engineers, it's extending AI-native internal tooling to the rest of Plaid. Tools built for engineers assume things non-engineers don't have: local toolchains, monorepos, engineer credentials, PR-based workflows. That mismatch means Ops, Support, and other teams can't easily inherit what we build for engineering. They need their own path and we're building that path. Role: As a Senior Software Engineer on Intelligent Tooling, you will build and operate internal systems that empower non engineering teams to automate their workflows with AI. You will own the product and platform layer for internal tools, including the constraints and infrastructure that keep those tools safe and maintainable. There's no existing playbook for this at Plaid. You will define what the right non-eng AI surface looks like, ship its first durable versions, and partner closely with internal users to make sure it solves real problems. You will act as the engineering point of contact embedded with non-engineering teams, running discovery and trans

O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

Technical Program Manager – Applied Infrastructure About the Team The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety. The Applied Infrastructure TPM team partners across engineering to lead foundational programs that ensure OpenAI’s infrastructure can meet current and future demand. About the Role We’re looking for a seasoned Technical Program Manager to drive critical infrastructure programs across the Applied organization. This TPM will focus on cross-cutting initiatives such as general compute capacity planning, process transformation, cost and quota attribution and optimization, and coordination across infrastructure and product stakeholders. There will also be focus on evolving OpenAI’s infrastructure to support growth, scale and new products. This work is core to how OpenAI manages and grows its infrastructure footprint in a disciplined, scalable way. Location: San Francisco, CA (Hybrid – 3 days/week in-office) In this role, you will: Serve as the DRI for complex infrastructure programs spanning CPU planning, orchestration, and other resource management domains (e.g. networking, storage). Build and operationalize systems to capture demand signals, model future capacity needs, and align infrastructure planning across internal teams and partners external to the company. Partner closely with Infrastructure, Product and Finance teams to forecast infrastructure usage patterns and ensure supply/demand alignment. Lead cost attribution and quota enforcement programs to promote stability and ensure equitable access to resources across teams. Drive simplification and standardization of infrastructure tooling and processes across Applied and Infra organizations. Drive cross functional programs to evolve our infrastructure to support new growth and scale Work with external v

AWSAzureRestAI
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT infrastructure team is responsible for ensuring that our products can serve rapidly growing demand with the performance, reliability, and quality our users expect. This work sits at the intersection of product demand, model deployment, inference, research, fleet, and capacity. The team translates changing product and model needs into clear capacity decisions and safe, scalable launches. About the Role We are seeking a Technical Program Manager to lead the operating system for Chat capacity and model deployment. You will connect demand forecasting and capacity allocation with model readiness, rollout planning, launch coordination, and post-deployment learning. You will also own mode deployment beyond capacity by working with cross functional teams across research, post-training, inference and product to own mainline model deployment. You will bring structure to constrained-capacity decisions, improve the tooling and mechanisms teams use to prioritize demand, and help new models reach users safely and efficiently. Success requires technical depth, sound judgment under ambiguity, and crisp execution across product, research, infrastructure, and operations teams. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own cross-functional programs for Chat capacity forecasting, allocation, headroom planning, and constrained-capacity operations. Build durable intake, prioritization, and decision mechanisms that connect product demand and model requirements to available serving capacity. Partner

AWSRestAIRust
O
📍 San Francisco, California, United States· Full-time
✓ High-confidence listingCompany trend -80.2%
Quick readStrong listing-quality and freshness signals

About the Team The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT engineering org builds and operates the systems that bring product improvements to users across backend services, web, mobile, and desktop platforms. The Developer Velocity team partners with product engineering, platform, infrastructure, reliability, engineering acceleration, and observability teams to make everyday development faster and releases safer, more predictable, and easier to operate. About the Role We are seeking a Technical Program Manager to improve developer velocity and deployment excellence across ChatGPT. You will lead durable improvements to local development, CI, testing, build systems, release trains, progressive rollout, and post-deployment validation. You will identify the highest-leverage sources of engineering friction, align teams around shared standards and metrics, and drive adoption of tooling and workflows that improve both speed and reliability. This role combines systems thinking, technical program leadership, and hands-on operating rigor across a broad engineering surface. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Own the cross-functional roadmap for improving local development, CI, testing, build workflows, and release infrastructure. Create a durable intake and prioritization mechanism for developer friction, using evidence to focus teams on the highest-impact improvements. Lead programs that improve deployment speed and safety, including pre-merge confidence, progressive rollout, release guardrails, rollback readiness, and post-deploy valida

AWSCI/CDRestAI
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and proactive Security Engineer to help us build, maintain, and continuously improve the security posture of our rapidly growing ML infrastructure platform. As one of the first dedicated security hires at Baseten, you will work cross-functionally with engineering and operations teams to ensure we’re meeting the highest standards of confidentiality, integrity, and availability. You’ll have an opportunity to shape our security strategy and best practices from the ground up, influencing the way our platform handles sensitive data for both internal and external stakeholders. RESPONSIBILITIES Security architecture and design: Collaborate with engineering teams to design and implement secure systems and infrastructure, including cloud (AWS/GCP) environments and container orchestration platforms. Vulnerability management: Lead proactive vulnerability assessments, pen tests, and remediation efforts to ensure our products and infrastructure remain secure. Incident response: Develop and maintain incident response processes, including detection, analysis, containment, eradication, and post-incident reviews. Identity and access management (IAM): Oversee IAM strategies and tools to ensure the right people have the right level of access to our systems and data. Security compliance and audits: Work closely with operations to ensure compliance with relevant standards (e.g., SOC 2, ISO 27001) and

AWSGCPCI/CDMachine Learning
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The ChatGPT Model Flywheel team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. Team Focus Areas Model Experimentation: Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management. Model Deployment: Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling. Automate capacity management and incorporate platform-wide health monitors. Model Measurement: Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards. Improve end-to-end feedback loops for continual model improvement. Key Partnerships Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API. In this role, you will: Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks. Drive expansion and improvement of multi-tier model experiences. Support and scale self-serve experiment capabilities and automated guardrails. Lead model rollout automation, capacity management, and health monitoring. Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.). You might thrive in this role if you have: Proven experience leading engineering teams in complex, cross-functional environments. Demonstrated success shipping production systems at scale (ideally for AI or large backend services). Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling. Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders. Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus. Why Work With Us Tackle highly impactful technical challenges at the cutting edg

AWSRestAIGo
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. About the Role OpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency of our systems. In this role, you'll apply deep technical expertise to optimize infrastructure and application-level performance across mission-critical products like ChatGPT and our developer API. You’ll work cross-functionally with teams building core services, training models, and developing real-time user experiences to push our latency, throughput, and cost-efficiency to the next level. We are looking for engineers who thrive in ambiguous environments, value deep systems understanding, and are motivated by delivering measurable impact. This is a highly technical, individual contributor role focused on root-cause analysis, profiling, instrumentation, and architecture-level performance improvements across our stack. In this role, you will: Analyze and optimize performance across application, middleware, runtime, and infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond. Develop tooling and metrics that provide deep observability into system performance. Collaborate closely with infra, platform, training, and product teams to identify key performance goals and drive systemic improvements. Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale. Lead investigations into high-impact performance regressions or scalability issues in production. Drive performance testing strategies and help define SLAs/SLOs around latency and throughput for critical systems. You might thrive in this role if you: Have 7+ years of experience in software engineering with a strong tr

PythonAWSRestAI
P
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -72.3%

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam. Making data driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide tooling and guidance to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. Engineers on Data Infrastructure are domain experts in Data Warehouse, Data Lakehouse, Spark, Workflow Orchestration, and Streaming technologies. We scale our existing data pipelines in a performant and cost efficient way while creating the necessary abstractions to make developing on top of this platform extremely simple for other engineers at Plaid. Responsibilities Contribute towards the long-term technical roadmap for data-driven and machine learning iteration at Plaid Leading key data infrastructure projects such as improving ML development golden paths, implementing offline streaming solutions for data freshness, building net new ETL pipeline infrastructure, and evolving data warehouse or data lakehouse capabilities. Working with stakeholders in other teams and functions to define technical roadmaps for key backe

PythonAWSMachine LearningAI
O
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -80.2%

About the Team The Ads Support Delivery team is responsible for helping successfully operate and grow on our Ads product. This includes technical guidance, troubleshooting complex delivery and monetization issues, and partnering closely with Product, Engineering, Trust & Safety and Go-To-Market teams to resolve customer-impacting problems and improve the platform over time. The team’s mission is to deliver a high-quality customer experience at scale by combining strong human support with automation, self-service, and AI-enabled workflows, while maintaining high operational rigor. About the Role: As a Support Delivery Lead for Ads, you will lead a team responsible for end-to-end support delivery across the ads ecosystem, including campaign setup, delivery, billing, measurement, and policy navigation. You will set the operational bar for quality, responsiveness, and consistency; coach and grow the team; and translate support signals into actionable improvements with Engineering, Product, and Go-To-Market partners. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: Lead and support a team of Ads support engineers, ensuring they have the tools, clarity, and coaching needed to operate at a high bar in a technically complex domain. Set clear expectations and operating standards, run recurring performance reviews, and build development plans that grow both technical depth (ad tech fluency) and customer-facing excellence. Design and continuously improve support coverage for ad buyers, ensuring the team can diagnose delivery issues and monetization and integration issues with equal rigor. Act as the bridge between Support Delivery, Engineering, Product, and Go-To-Market teams. Drive alignment on priorities, escalation paths, launch readiness, tooling improvements and mechanisms to reduce repeated customer pain points. Partner with engineering teams

AWSRestAIGo
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We are seeking an experienced and detail-oriented GRC (Governance, Risk, and Compliance) Manager to build, support, and continuously enhance Baseten’s security governance, compliance, and privacy programs. As one of the early members of our security organization, you will play a key role in ensuring our platform meets and exceeds the highest standards for privacy, trust, and regulatory compliance. In this role, you’ll work cross-functionally with engineering, operations, legal, and leadership teams to develop policies, manage audits, and implement controls aligned with frameworks such as SOC 2, ISO 27001, ISO 27701, and FedRAMP. You’ll be instrumental in building scalable processes to manage risk, support customer assurance, and uphold Baseten’s commitment to security and compliance as we grow. RESPONSIBILITIES Governance & Policy Development: Design, implement, and maintain security governance frameworks, policies, and procedures that align with Baseten’s risk posture and industry best practices. Risk Management: Build and manage the company-wide risk assessment program, identifying, tracking, and mitigating key security and compliance risks. Compliance Operations: Lead efforts to achieve and maintain compliance with SOC 2, ISO 27001/27701, HIPAA, FedRAMP and other applicable standards and regulations. Audit & Certification Management: Coordinate external audits and certification processes, ensuring e

AWSGCPMachine LearningAI
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Manager focused on TPUs at Baseten, you will lead the "engine room" for our non-NVIDIA accelerator fleet, architecting, securing, and optimizing the Google Cloud TPU (and broader emerging accelerator) capacity that powers our customers' AI workloads. You'll own the end-to-end journey of capacity management for this fleet, from securing large-scale TPU pod allocations to building the automation that ensures reliable uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering, with a specific focus on the TPU ecosystem. You will act as the fleet orchestrator for Google's TPU architecture, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics as we diversify beyond NVIDIA. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the latest generation of TPU hardware, like Google's Trillium (v6e) architecture, and partnering closely with the Model Performance (MP) team to ensure workloads are tuned for TPU-specific execution. EXAMPLE INITIATIVES The TPU Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's TPU clusters, including pod slicing and topology planning Global Workload Orchestration: Bui

PythonAWSAzureGCP
B
📍 San Francisco, California, United States· Full-time
✓ Quality checkedCompany trend -79.1%

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer at on the Training Infrastructure team, you'll architect and lead development of our training platform, supporting top tier research engineers and model developers. You'll make key technical decisions for the infrastructure enabling developers to deploy, scale, and monitor their workloads with high performance and reliability. You’ll own scheduling, storage, networking, reliability, and observability of technical systems in the training stack EXAMPLE INITIATIVES Take a look at what we’ve built so far: Overview of the product so far Training docs overview Story of the Training product Research we've done RESPONSIBILITIES Design and architect scalable infrastructure systems for our ML training platform (e.g. scheduling, storage, and networking) Partner closely with developers and research engineers to translate complex training requirements into technical solutions Design and architect a global training scheduler Design and architect reinforcement learning systems and continuous learning pipelines Drive long-term improvements to improve reliability of systems and velocity of development Partner closely with SRE and Capacity teams to unlock state of the art training infrastructure Make critical architectural decisions balancing performance with system reliability Lead technical discussions and mentor junior engineers on infrastructure best practices Contribute to long-term technical strateg

PythonAWSGCPKubernetes
B
📍 San Francisco, California, United States· Full-time· Remote
✓ High-confidence listingCompany trend -79.1%
Quick readStrong listing-quality and freshness signals

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F , led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products. THE ROLE As a Global Capacity Lead at Baseten, you will lead the "engine room" of the company, architecting, securing, and optimizing the global GPU fleet that powers our customers' AI workloads. You’ll own the end-to-end journey of capacity management, from securing multi-million dollar GPU clusters to building the automation that ensures 99.9% uptime across multi-cloud environments. This role is a great fit for entrepreneurial engineers who want to bridge the gap between high-finance asset management and deep infrastructure engineering. You will act as the fleet orchestrator for the world's most advanced chips, ensuring Baseten never experiences a capacity outage while maintaining elite unit economics. To be clear, this is a high-stakes engineering role. You will be hands-on with Kubernetes orchestration while also leading specialized pods focused on the next generation of hardware, like NVIDIA’s Blackwell (B200) architecture. EXAMPLE INITIATIVES The B200 Frontier: Architecting the infrastructure readiness and deployment strategy for Baseten's first Blackwell GPU clusters. Global Workload Orchestration: Building "Multi-cloud Capacity Management" systems to move customer workloads seamlessly across regions to optimize cost and latency. Precision GPU Triage: Developing automated Go-based operators to identify, cordon, and repair unhealthy H100 nodes in under an hour. The Supply Chain of Intelligence: Partnering with lead

PythonAWSAzureGCP
🔔

Get new aws and tooling platform lead jobs in San Francisco, United States by email

Daily job updates · Unsubscribe anytime